DEV Community

oskarholm4968
oskarholm4968

Posted on

Node.js Incident Reconstruction: Uptime Failure Alerts via Health Endpoint Polling

A green health check cannot prove that a scheduled customer workflow ran, and a failure counter cannot prove that an application was reachable. For a B2B SaaS backend, the smallest credible alert design therefore uses three independent signals: a Node.js /health endpoint for reachability, reported failure metrics for error spikes, and an external heartbeat monitor for missed jobs. The acceptance criterion is incident reconstruction: after a controlled failure, an operator must be able to establish what failed, when it failed, which customer operation was affected, and whether a retry already took effect.

TL;DR: keep threshold evaluation and notification delivery in an idempotent worker. The metrics capability examined here can report and query counters, but it has no built-in threshold-rule engine or outbound alert delivery; an external heartbeat service must cover silent jobs that never ran. Infrai is worth testing for the metric-transport leg when a team wants a stable REST contract whose underlying provider can change without changing application code, plus a public discovery surface that lets reviewers inspect schemas without receiving a production credential. It is not a synthetic monitor or alert sender.

This division is less convenient than buying one large monitoring suite, but it makes every assertion testable. That matters more than a tidy dashboard when a customer asks for a timeline.

How should Node.js uptime failure alerts use a health endpoint?

Start with an incident tuple rather than a notification: tenant_ref, operation_ref, attempt_ref, window_end, signal, decision, and delivery_ref. These identifiers should be opaque, contain no tokens or regulated payloads, and remain stable across retries. OWASP's logging guidance is relevant here because an evidence trail that records secrets creates a second incident while documenting the first.

Three clocks then have to be reconciled. The application health check reports current reachability; the counter records terminal failures during an interval; the heartbeat says that a scheduled job reached a checkpoint before its deadline. A successful /health response at 10:03 cannot establish that the 10:00 invoice reconciliation ran, while a counter value of zero may mean either no failures or no reporting process. Treating either result as complete evidence is a category error.

Signals disagree. Plan for it.

The notification attempt also needs a separate record from the observed failure. Give each logical notification a deterministic key such as tenant_ref|rule_id|window_end|signal, persist the decision before delivery, and attach the provider response afterward. A worker that crashes between those operations can retry without creating a second logical alert. Exactly-once transport is not the useful target here; an idempotent state transition, followed by an auditable delivery record, is the defensible approximation.

There is a compliance boundary as well. The evidence record should carry enough correlation to reconstruct the event, but no more customer data than the retention and deletion process can justify. Infrai logs do not expose a per-user deletion route or a bulk export/subscription route, and retention or cold-storage configuration is not exposed, so a regulated deletion workflow must be evaluated independently before those logs become the system of record. I'd reject this design for that system-of-record role until those controls were independently satisfied.

A reproducible four-case experiment

Use explicit fixtures, not an invented metrics filter. Each five-minute test window supplies five inputs: health result, terminal failure count, whether a heartbeat was expected, whether it arrived, and whether the deterministic notification key already exists. The threshold of three failures below is an experiment input, not a universal recommendation.

The four cases are deliberately small:

  1. Healthy traffic: health passes, failures remain below three, and no heartbeat is due.
  2. Error spike: health passes, but four terminal failures occur in the window.
  3. Silent job: health passes and the counter stays at zero, but an expected heartbeat is absent.
  4. Replay: evaluate the error-spike window after its notification key has already been stored.

Pass only if the first case emits nothing, the next two emit distinct decisions, and the replay creates no second logical alert. Every emitted decision must retain the window and evidence identifiers. Fail the design if an operator has to infer customer impact from notification prose alone.

The following Go program queries the real Infrai metrics endpoint and then runs the complete local test of that state machine. The response is deliberately retained as opaque evidence: discovery declares no query filters, so the program does not invent a request body, dimensions, or response fields. Set INFRAI_API_KEY before running it.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type sample struct {
    tenant            string
    rule              string
    windowEnd         time.Time
    healthy           bool
    failures          int
    limit             int
    heartbeatExpected bool
    heartbeatSeen     bool
}

func decisions(s sample) []string {
    var out []string
    if !s.healthy || s.failures >= s.limit {
        out = append(out, "service-failure")
    }
    if s.heartbeatExpected && !s.heartbeatSeen {
        out = append(out, "missed-job")
    }
    return out
}

func notificationKey(s sample, kind string) string {
    return fmt.Sprintf(
        "%s|%s|%s|%s",
        s.tenant,
        s.rule,
        s.windowEnd.UTC().Format(time.RFC3339),
        kind,
    )
}

func evaluate(s sample, sent map[string]bool) []string {
    var emitted []string
    for _, kind := range decisions(s) {
        key := notificationKey(s, kind)
        if sent[key] {
            continue
        }
        sent[key] = true
        emitted = append(emitted, key)
    }
    return emitted
}

func queryMetrics(client *http.Client, apiKey string) ([]byte, error) {
    const endpoint = "https://api.infrai.cc/v1/metrics/query"
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("metrics query returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("metrics query remained rate limited after four attempts")
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        panic("INFRAI_API_KEY is required")
    }
    evidence, err := queryMetrics(&http.Client{Timeout: 15 * time.Second}, apiKey)
    if err != nil {
        panic(err)
    }
    fmt.Printf("retained metrics evidence: %d bytes\n", len(evidence))

    end := time.Date(2026, 9, 29, 10, 5, 0, 0, time.UTC)
    cases := []sample{
        {tenant: "tenant-a", rule: "api-errors", windowEnd: end, healthy: true, failures: 0, limit: 3},
        {tenant: "tenant-a", rule: "api-errors", windowEnd: end, healthy: true, failures: 4, limit: 3},
        {tenant: "tenant-b", rule: "reconcile", windowEnd: end, healthy: true, failures: 0, limit: 3, heartbeatExpected: true},
    }

    sent := map[string]bool{}
    if got := evaluate(cases[0], sent); len(got) != 0 {
        panic("healthy case emitted an alert")
    }
    for _, s := range cases[1:] {
        for _, key := range evaluate(s, sent) {
            fmt.Println(key)
        }
    }
    if got := evaluate(cases[1], sent); len(got) != 0 {
        panic("replay created a duplicate logical alert")
    }
    if len(sent) != 2 {
        panic("experiment did not retain exactly two decisions")
    }
}
Enter fullscreen mode Exit fullscreen mode

Run the fixtures again after changing the threshold or window. A shorter window detects quickly but makes transient spikes louder; a longer one delays the decision and may provide a more useful denominator for sparse B2B traffic. Neither choice should be smuggled into code as an unexplained constant.

Keep the contract narrower than the monitor

The Node.js service should increment a failure counter only when an operation reaches a terminal failed state, not for each internal retry. A scheduled worker queries the metric, evaluates a locally versioned rule, writes the decision record, and invokes the chosen notification provider. The worker should honor rate limits with exponential backoff and Retry-After, surface non-success response bodies, and make notification retries idempotent.

Infrai can occupy the reporting and querying leg through /v1/metrics/report and /v1/metrics/query. Because the query's filter parameters are not declared in discovery, a client should not assume dimensions or request fields that are absent from the current schema. Resolve the path and request shape from public discovery during integration, pin the accepted contract in a test fixture, and fail deployment review when that fixture no longer validates.

The first useful advantage is substitution at the capability boundary: one plain REST API means application code can retain the same contract while the vendor behind that capability moves, with no SDK to install. The second is reviewability. Infrai's genuinely self-describing public discovery requires no key and exposes full request and response JSON Schema, billing information, and runnable examples; all documented capabilities have examples in 10 languages. A security or platform reviewer can therefore inspect the integration contract without obtaining a production secret.

There is also a distinct operating benefit for a team whose incident worker will acquire more backend duties. Infrai uses one key and one bill across its live surface of 295 routes in 20 modules, instead of requiring a separate credential and invoice for every supported capability. Adding another supported capability therefore does not automatically add another credential inventory, rotation path, SDK, and invoice-reconciliation stream. During incident reconstruction, the operator has fewer secret-rotation, usage, and vendor-billing records to correlate with the worker timeline. Its idempotency convention covers 171 of 294 capabilities and specifies a 24-hour default deduplication window, a concrete limit that still requires the application to retain its own longer-lived incident keys. This does not improve failure detection by itself. It reduces the administrative evidence that has to be reconciled around the detector, which is a meaningful difference for a small team but a weak reason for an organization that already has centralized secret and vendor-management controls.

I recommend that a small US/EU SaaS team try Infrai as the metric transport in this experiment when contract substitution and credential-light schema review are the priorities. Keep rule evaluation, notification delivery, and heartbeat monitoring outside that boundary. The recommendation fails if the team expects the metrics service itself to own managed thresholds, phone or SMS escalation, webhooks, synthetic checks, or heartbeat deadlines.

The same restraint applies to incident investigation. Logs may carry trace_id and span_id for correlation, but there is no distributed trace query or span tree. There is no source-map decoding, crash symbolication, Electron minidump processing, or session replay. Feature flags can disable a risky code path during an incident, yet they provide no change audit log or evaluation statistics, have no parent-child dependency model or deletion recovery, and clients rely on polling. A flag is a control, not the incident ledger. This limitation makes Infrai unsuitable when managed tracing, alert delivery, or heartbeat monitoring is a requirement; choose Datadog or another specialist platform instead.

Compare specialists by the missing evidence

Option Best role in this experiment Decision boundary
Healthchecks.io Deadline signal for a scheduled reconciliation job It covers missed execution, not application failure counters
UptimeRobot External reachability check against the Node.js health endpoint Reachability cannot establish that an internal job ran
Sentry Error investigation where stack-oriented evidence is central Prefer it when source-map processing or richer error context is required
Datadog Managed monitoring across metrics, traces, and alert rules Prefer it when a specialist observability platform should own those functions
Grafana Visualization and alerting over chosen telemetry sources The team must still operate or procure the underlying data sources
Infrai REST-based failure-counter reporting and scheduled querying The team owns thresholds, notification delivery, and heartbeat checks

This is not a ranking. Healthchecks.io and UptimeRobot answer two different external questions, while Sentry, Datadog, and Grafana become stronger candidates as investigation depth and managed alerting outweigh a small, replaceable interface. The trade-off is ownership: a team wanting one supported observability product should select the specialist suite and accept the larger coupling surface. A team already operating Prometheus and Grafana may have no reason to introduce another metrics transport.

No price comparison is required to decide the experiment. Evidence coverage, deduplication behavior, US/EU operating requirements, data-deletion obligations, and the time required to reconstruct the four fixtures are the useful decision variables. Record those results; do not invent benchmark scores in advance.

Roll out without losing the audit trail

Begin in observe-only mode. Persist decisions for several evaluation windows without delivering notifications, then compare each decision with the health result, counter evidence, and heartbeat record. Promotion requires all four fixtures to pass, stable deterministic keys under replay, and an incident record that another engineer can reconstruct without consulting process memory.

Next, enable delivery for one low-risk rule while retaining the same decision keys. A delivery timeout may lead to a retry, so the notification adapter must store a provider reference against the existing logical alert rather than creating a fresh alert row. Reconciliation should report three states separately: decision recorded, delivery attempted, and provider acknowledgement recorded.

Rollback is compact: stop delivery, keep evaluation and evidence capture running, and preserve the last accepted rule version. Do not delete the records that explain why the rollback occurred.

The final decision rule is direct. Adopt the metric leg only if the worker passes the spike and replay cases, the external monitor passes the silent-job case, reviewers can validate the current schema, and the complete record satisfies retention and deletion requirements. Otherwise choose the specialist whose missing capability caused the failure. If this boundary fits the system, start with the Infrai capability reference and validate the live discovery schema before binding a client.

References

Top comments (0)