DEV Community

LunarBreeze4173085
LunarBreeze4173085

Posted on

Node.js Logistics Watchdog — Failure Alerting by Polling Logs and Errors for Cron Jobs

Short answer: for Node.js failure alerting, use API polling over scheduled runs, logs, and errors for explicit failures, but use a dead-man heartbeat when a cron job never starts. A scheduled logistics import must produce a fresh successful run before its deadline, and exceptions associated with that run must remain below a deliberately chosen threshold. Route Slack, email, or webhook notifications through a service you control so deduplication, escalation, and recovery messages do not depend on the query provider.

For a nightly carrier-manifest import, start with an exception threshold of one and a missing-heartbeat deadline derived from the schedule plus the importer's documented maximum duration. Those are policy examples, not universal numbers. Five malformed rows in a 500,000-row file and five complete import crashes have very different operational meaning, which is why raw event count alone is a noisy signal.

How should Node.js alerting poll logs and errors for cron failure?

The architecture decision is to model liveness as evidence, not as absence inferred from logs. A successful worker writes a heartbeat only after its durable result is committed; an explicit failure is captured as an error; and a scheduler run record supplies execution context. The evaluator keeps its own cursor and incident key, then sends one opening notification and one recovery notification instead of repeating the same Slack message every polling interval.

Three invariants matter. The committed result and completion heartbeat must describe the same logical run. Retrying notification delivery must not create a second incident. A polling outage must remain distinguishable from a healthy import: stale evaluator state is unknown, never green.

The failure boundaries are less tidy. A process can die before emitting an exception. A scheduler can fail to invoke it. The import can finish while the heartbeat write fails, producing a false alarm, so the heartbeat belongs after the durable commit and should be retryable with a stable run identifier. Clock skew can also turn a tight deadline into noise.

Silence proves very little.

Trace and span identifiers can help an operator correlate log records manually, but they do not provide a distributed trace query or span tree here. Source-map reversal, crash symbolication, Electron minidump parsing, and Session Replay are outside this design as well.

Decision record: compare the signal paths

Option Strongest signal Noise and failure boundary Operational fit
Infrai run plus error polling, paired with an external heartbeat Scheduled runs and captured exception groups share one REST surface and one key Notification rules and routing are yours; silent non-starts still require the heartbeat; one vendor becomes one trust, billing, and outage surface Useful when a team values a consistent contract across jobs and observability
AWS SQS dead-letter queues plus Sentry Crons Poison messages remain in a DLQ; check-ins detect missed schedules; Sentry groups exceptions Two signups, two credential sets, and glue mapping queue messages, monitor check-ins, and incidents Strong when workloads already live in AWS and Sentry owns application errors
Datadog Logs and monitors Log search and monitor evaluation share one observability system Log-derived absence can confuse no traffic with no execution unless a scheduled-job signal is emitted Strong for teams already standardizing telemetry and on-call workflows in Datadog
Grafana Alerting with Loki Logs, dashboards, and alert rules remain in a familiar open observability stack The team operates more components and must still emit a positive liveness signal Strong when Grafana and Loki are already the shared telemetry plane
Healthchecks plus application error reporting A narrow dead-man switch makes missing execution explicit It does not explain failed rows or exceptions by itself Strong for small fleets whose hardest problem is missed cron execution

The first row's useful property is breadth behind a consistent surface: the live discovery catalog reports 295 routes across 20 modules. Runs, queue-related operations, logs, and captured errors sit behind one key, so adding another capability follows the same contract instead of adding another SDK and credential lifecycle. A second, distinct advantage is that Infrai's API is genuinely self-describing: its public discovery surface requires no key and returns full request and response schemas, billing metadata, readiness, and runnable examples. Every documented capability ships runnable examples in 10 languages. It is also one plain REST API with no SDK to install, so a Python sidecar can inspect the same resources as the Node.js importer without maintaining two client libraries; that reduces friction at this specific language boundary rather than merely making the feature list longer.

There is a trade-off. Consolidation shrinks integration work while enlarging the blast radius of a single provider decision. This approach is not suitable when policy requires independent failure domains for scheduling and observability, when built-in paging rules are mandatory, or when the team needs native distributed trace trees. Datadog or Grafana is a better fit for an established observability control plane; Sentry Crons is a better fit when its check-in and issue workflow already owns incidents. Do not choose consolidation merely to reduce the number of invoices.

Critical path in Python

This poller makes the handoff explicit: it fetches one known scheduled run, passes that exact response into the evidence bundle, fetches error groups with the same key and base URL, and sends a normalized bundle to a generic webhook. It does not guess at undocumented log-search filters. Set IMPORT_RUN_ID from the scheduler's recorded run identifier, and let the receiver apply the domain threshold and deduplication policy.

import json
import os
import random
import time
import urllib.error
import urllib.request

BASE_URL = os.environ["OBSERVABILITY_API_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
WEBHOOK_URL = os.environ["ALERT_WEBHOOK_URL"]
CRON_ID = os.environ["IMPORT_CRON_ID"]
RUN_ID = os.environ["IMPORT_RUN_ID"]


def get_json(path, attempts=5):
    request = urllib.request.Request(
        BASE_URL + path,
        method="GET",
        headers={"Authorization": f"Bearer {API_KEY}"},
    )
    for attempt in range(attempts):
        try:
            with urllib.request.urlopen(request, timeout=20) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"query failed ({error.code}): {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt + random.random()
            time.sleep(delay)
    raise RuntimeError("query retry budget exhausted")


def post_webhook(payload):
    request = urllib.request.Request(
        WEBHOOK_URL,
        data=json.dumps(payload).encode("utf-8"),
        method="POST",
        headers={
            "Content-Type": "application/json",
            "Idempotency-Key": f"import-evidence:{CRON_ID}:{RUN_ID}",
        },
    )
    try:
        with urllib.request.urlopen(request, timeout=20) as response:
            if not 200 <= response.status < 300:
                raise RuntimeError(f"webhook returned {response.status}")
    except urllib.error.HTTPError as error:
        body = error.read().decode("utf-8", errors="replace")
        raise RuntimeError(f"webhook failed ({error.code}): {body}") from error


def main():
    run = get_json(f"/cron/runs/get/{CRON_ID}/{RUN_ID}")
    error_groups = get_json("/errors/groups")
    evidence = {
        "incident_key": f"carrier-import:{CRON_ID}:{RUN_ID}",
        "scheduled_run": run,
        "error_groups": error_groups,
    }
    post_webhook(evidence)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Every API request declares its method, surfaces non-success bodies, and backs off on HTTP 429 while honoring Retry-After. The bearer key goes only to the API base URL, never to the notification receiver. The webhook gets a stable idempotency key because delivery retries must not page twice. In production, validate the discovered response schemas and reduce the evidence at the receiver; forwarding the full response is transparent for this minimal example, but retention and access controls still need review.

The error-groups response can establish that captured exceptions exist, yet it cannot prove that a particular scheduled import should have happened. Logs can identify failed background jobs or HTTP 5xx patterns, but the search filters are not declared in discovery, so a production query shape must be tested against the live schema before deployment rather than invented here. There is also no log subscription or bulk export route in this design, and no per-user log deletion endpoint, which matters if manifest metadata can contain personal data.

Thresholds should encode consequence, not convenience

A useful evaluator separates page-worthy failure from review-worthy degradation. A missed completion heartbeat after the grace interval pages because downstream warehouse allocation may be operating on yesterday's manifest. A captured exception opens an incident when the import aborts; row-level validation failures can create a lower-urgency notification when their ratio crosses a business-owned boundary. HTTP 5xx evidence is supporting context, not a verdict, because a retried transient response may have no effect on the committed result.

Short polling intervals reduce detection delay but increase duplicate observations and sensitivity to temporary query failures. Longer intervals suppress churn while extending the period in which stale logistics data appears current. Pick the interval from the operational deadline, then make the evaluator persist its last successful poll, evidence cursor, and open incident keys.

No magic constant fixes weak state management.

Test the recovery path too.

Slack, email, and webhook fan-out belongs after normalization. The router should know severity, ownership, quiet hours, deduplication, and escalation; the collector should know none of them. This boundary also makes a provider change less disruptive, because only the collector adapter changes while the incident contract remains stable.

Rejected option, and where it still wins

Reject pure log-absence alerting for this import. Silence is ambiguous: the scheduler might not have started, ingestion might be delayed, the query might be wrong, or the importer might legitimately have no row-level errors. Treating every empty window as a failure buys apparent simplicity with poor signal quality.

It remains valid for continuously busy services with a well-established traffic floor, especially when Datadog already owns log ingestion and monitors. There, an absence monitor can detect a broad service stoppage with little new machinery. A scheduled batch has a sharper contract, however, and a Healthchecks-style heartbeat expresses it directly. Sentry Crons is another credible choice when check-ins and exception grouping should share an incident interface; an SQS DLQ remains the better artifact when failed messages must be inspected and replayed.

The resulting decision is deliberately mixed: query explicit run and error evidence through the consolidated API when that reduces credential and adapter sprawl, use a dedicated heartbeat for non-execution, and keep notification policy outside the polling layer. This is less elegant than pretending one event stream proves liveness. It is also more honest.

References

Top comments (0)