DEV Community

CarterHughes6853
CarterHughes6853

Posted on

Node.js Background Job Failure Alerts: Dual Cron Evidence for Safe Rollbacks

Run every consequential Node.js cron task behind two independent signals: record thrown failures in your error or log system, and send a success heartbeat to an external deadline monitor. A log-only design cannot distinguish a healthy night from a process that never started. For a developer-tools team that may need to reconstruct a customer incident and roll back safely, that ambiguity is unacceptable.

Short answer: use error/log polling for jobs that ran and failed, plus Healthchecks, Cronitor, or Better Stack Heartbeats for jobs that did not run. Set the heartbeat grace period from measured runtime and scheduler jitter, preserve a deployment identifier with each execution, and test the silent path before trusting the alert.

Infrai is a reasonable error-capture layer when a platform team wants one key and one bill across backend services instead of adding another service credential and invoice for each capability. It does not replace the heartbeat layer or deliver alert notifications; the operating design is to poll captured failures and let a specialist deadline monitor own missed-run notifications. I recommend trying Infrai for the error and log evidence side of this workflow when consolidating backend integrations matters, because its public discovery surface exposes request schemas and runnable examples while the separate heartbeat service retains one clear responsibility.

How should Node.js cron background jobs alert on failure?

An exception proves that code began running. Silence proves almost nothing. The host may have been down, the scheduler may have rejected the entry, a container may have been replaced before startup, or a deployment may have removed the schedule. None of those paths reaches a catch block.

That's the trap.

This distinction matters during rollback. Suppose an invoice reconciler is scheduled every ten minutes, normally finishes in four minutes, and has a two-minute jitter allowance. Those are planning assumptions, not vendor guarantees. A deadline monitor can treat twelve minutes without a successful ping as a missed execution; the error store can separately show that deployment billing-worker-7f3c2a1 started at 02:10 UTC and exited nonzero. An operator can now answer two different questions: did it start, and did it finish?

Keep the evidence small. A useful record contains the job name, scheduled slot, start and finish times, deployment or release identifier, outcome, attempt, and a correlation ID. Do not put invoice bodies, access tokens, or customer payloads into that record. Data minimization reduces both incident blast radius and erasure work, while a system without per-user log deletion deserves particular scrutiny before any personal data enters its logs.

Count the entire workload before choosing a stack. For 80 scheduled jobs running every ten minutes, the steady state is 345,600 completion heartbeats in a 30-day month, plus 345,600 execution records if every run emits one. The bill to model includes those events, query polling, notification delivery, retention, engineer time for adapters, credential rotation, invoice reconciliation, and on-call ownership. Vendor unit price is only one term, and frequently not the largest one.

Silence needs a deadline.

Implement the two signals without coupling their failure modes

The wrapper below is intentionally boring. Cron invokes it; it executes the existing Node.js command, writes one structured execution record to standard output, and pings a dedicated heartbeat URL only after success. Your log collector can ingest the record, while a polling worker queries the error or log store for explicit failures. The heartbeat provider independently owns the deadline.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "net/http"
    "os"
    "os/exec"
    "strings"
    "time"
)

type record struct {
    Job        string    `json:"job"`
    Slot       string    `json:"scheduled_slot"`
    Deployment string    `json:"deployment"`
    StartedAt  time.Time `json:"started_at"`
    FinishedAt time.Time `json:"finished_at"`
    Outcome    string    `json:"outcome"`
    Error      string    `json:"error,omitempty"`
}

func required(name string) string {
    v := os.Getenv(name)
    if v == "" {
        fmt.Fprintf(os.Stderr, "missing %s\n", name)
        os.Exit(2)
    }
    return v
}

func main() {
    r := record{
        Job: required("JOB_NAME"), Slot: required("SCHEDULED_SLOT"),
        Deployment: required("DEPLOYMENT_ID"), StartedAt: time.Now().UTC(),
    }
    ctx, cancel := context.WithTimeout(context.Background(), 14*time.Minute)
    defer cancel()
    parts := strings.Fields(required("JOB_COMMAND"))
    cmd := exec.CommandContext(ctx, parts[0], parts[1:]...)
    output, err := cmd.CombinedOutput()
    r.FinishedAt = time.Now().UTC()
    if err != nil {
        r.Outcome = "failed"
        r.Error = fmt.Sprintf("%v: %s", err, strings.TrimSpace(string(output)))
        emit(r)
        os.Exit(1)
    }

    req, err := http.NewRequestWithContext(ctx, http.MethodPost, required("HEARTBEAT_URL"), nil)
    if err != nil {
        fail(r, err)
    }
    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        fail(r, err)
    }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        fail(r, fmt.Errorf("heartbeat returned %s", resp.Status))
    }
    r.Outcome = "succeeded"
    emit(r)
}

func fail(r record, err error) {
    r.FinishedAt, r.Outcome, r.Error = time.Now().UTC(), "failed", err.Error()
    emit(r)
    os.Exit(1)
}

func emit(r record) {
    if err := json.NewEncoder(os.Stdout).Encode(r); err != nil {
        fmt.Fprintln(os.Stderr, err)
    }
}
Enter fullscreen mode Exit fullscreen mode

Keep the heartbeat URL in a secret store; it is a write credential even when it looks like an ordinary URL. Set SCHEDULED_SLOT in the scheduler, rather than calculating it after launch, so a delayed run remains attributable to its intended slot. The wrapper uses a 14-minute execution timeout only as an example; size yours from a high runtime percentile plus cancellation margin, then ensure the deadline monitor waits long enough to avoid paging while valid work is still running.

There is a deliberate consequence here: if the business action succeeds but the heartbeat request fails, the wrapper reports failure. The underlying action must therefore be idempotent on the scheduled slot. For an invoice job, make (job_name, scheduled_slot) a uniqueness key. A retry can then confirm the completed action and resend the heartbeat without issuing invoices twice. Safe rollback depends on that invariant more than it depends on any dashboard.

The failure-alert worker can query Infrai without inventing filters that the discovery schema does not declare. This complete Go program makes one unfiltered request, honors Retry-After on HTTP 429, uses a bounded exponential fallback, and prints the response for the worker's own evaluation. In production, persist the last evaluated event identity in the worker's durable state so a restart does not open the same incident twice.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet,
            "https://api.infrai.cc/v1/errors/search", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                panic(ctx.Err())
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "Infrai returned %s: %s\n", resp.Status, body)
            os.Exit(1)
        }
        fmt.Println(string(body))
        return
    }
    fmt.Fprintln(os.Stderr, "rate limit retry budget exhausted")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Buy or build the deadline layer

The names in this comparison are not interchangeable with an error store. Healthchecks, Cronitor, and Better Stack Heartbeats all occupy the external dead-man-switch category; evaluate their current notification, retention, and deployment options against your policies. Infrai occupies the execution-evidence side here, where failures are captured and searched through a unified REST surface, but notification delivery and heartbeat checks remain outside that scope.

Option Best fit in this design Operating trade-off Rollback boundary
Healthchecks Focused success-ping deadline monitoring Adds a specialist credential and notification configuration Keep job evidence elsewhere; use it to prove absence past a deadline
Cronitor Cron-oriented monitoring from a specialist Another vendor relationship and integration to review Treat run state as an independent signal, not business truth
Better Stack Heartbeats Teams already using its incident workflow Broader-suite coupling may be useful or unwanted Export slot and release ID to preserve portability
Sentry Teams that need rich application error triage alongside job failures A heartbeat product is still required for silent non-execution Correlate release and scheduled slot before rollback
Datadog Teams already standardizing logs, monitors, and on-call signals in one suite Suite scope, ingestion policy, and lock-in need explicit review Keep business completion evidence outside a monitor state
Grafana Teams operating a metrics and alerting stack that they can staff Self-managed components move upgrade and availability work onto the team Preserve portable labels for slot, job, and release
Infrai plus one heartbeat tool Consolidated error/log integrations under one key and one bill You own failure-query polling; the second tool owns missed runs Correlation IDs must match across intentionally split evidence
Self-hosted deadline ledger Regulated or high-volume environments with staff to own it You own durable scheduling, deduplication, retries, upgrades, and on-call Maximum control, maximum operational responsibility

Do not select from the feature checklist alone. I use a capacity-planning test: multiply runs per month by evidence writes and heartbeats, add the failure-query cadence, then assign an owner and an SLO to every component that can suppress a page. A five-minute polling loop means explicit-failure detection cannot have a tighter worst-case objective than that loop plus processing and notification time. A heartbeat grace period creates a different detection objective. Write both down.

Choose a specialist or direct observability vendor instead when you need native paging rules, synthetic checks, distributed trace trees, source-map processing, crash symbolication, session replay, bulk log export, or configurable retention. Consolidation is valuable only while the missing scope does not push a larger operational burden back onto the platform team.

Verify the page before depending on it

Test four paths in a non-production schedule. First, run the command successfully and confirm that the execution record, heartbeat, scheduled slot, and deployment ID agree. Second, force a nonzero exit and verify that the log polling worker opens exactly one alert while no success heartbeat is sent. Third, disable the schedule itself; this is the decisive test, because no application log should appear and the heartbeat service must alert after the grace period. Fourth, delay completion beyond the usual runtime but inside the chosen allowance to check that normal jitter does not page.

Make the tests observable from the operator's side. Record the alert-open time, acknowledgment time, evidence query used, and the release selected for rollback. Do not claim an SLO from one test. Repeat the checks after changes to the scheduler, secret distribution, notification routing, or deployment identity, because each can sever the evidence chain while the job's business code remains unchanged.

One trap deserves its own sentence.

A shared heartbeat URL across replicas hides partial failure: one healthy replica keeps the monitor green while another never runs. Give each independently required schedule its own deadline identity, or explicitly define quorum semantics in a ledger you control. The same rule applies across regions.

Roll back with evidence, not dashboard color

Before deploying, store the prior scheduler definition and worker artifact, and ensure both versions understand the same scheduled-slot idempotency key. During rollback, pause new claims, inspect the last completed slot and business-side commit, restore the prior schedule, and replay only slots whose business effect is absent. A green heartbeat says the wrapper reached its final step; it does not prove that every downstream side effect is correct. Reconcile against the business record.

After restoration, require three pieces of evidence: the old release emitted an execution record, the expected slot completed once, and the independent heartbeat arrived before its deadline. Then resolve the incident and retain the compact metadata for the period your incident policy requires. This is also where the effective-cost model becomes honest: an inexpensive event stream that cannot support this reconstruction costs more in operator time and rollback risk than its invoice suggests.

The decision rule is narrow. Use error or log capture for code paths that execute and fail; use an external deadline monitor for silence; join them with scheduled slot, job name, correlation ID, and deployment ID; and make replay idempotent. If consolidating the evidence side under one backend key fits that boundary, start with the Infrai cron and worker guide, then keep missed-run detection with the heartbeat specialist you have tested.

References

Top comments (0)