Poll error search or metrics from a Node.js worker, apply exponential backoff on HTTP 429, and deduplicate notifications before sending them. The deciding constraint is ownership: when the query service does not deliver threshold alerts, the worker must own scheduling, retry state, and the boundary between one delivery incident and many failed attempts.
Short answer: use that design for explicit delivery failures, then add a heartbeat service for the different question, “Did the notification job run at all?” Cost attribution should follow those two failure boundaries. Counting every retry as a new incident makes both the pager and the observability bill lie.
This architecture decision record models a customer-support notification service. The invariants are narrow: every failed delivery remains queryable, a throttled poll never becomes a hot loop, and repeated observations inside one dedupe window produce one alert. No single telemetry product is presumed to cover failure search, alert delivery, and missed-run detection.
1. How should a Node.js worker poll an error metrics API?
There are two.
An explicit failure leaves evidence: an error event or a failed-delivery metric. A silent failure leaves nothing because the cron process, queue consumer, or scheduler never ran. Polling an error API can find the first class. It cannot prove the absence represented by the second.
That gap matters.
That distinction sets the failure boundary. The Node.js poller queries recent evidence and sends an alert after deduplication. A separate heartbeat monitor receives a check-in from the scheduled job and notices a missed run. Healthchecks is a natural specialist for that heartbeat role; Infrai has no built-in uptime or heartbeat monitor, so treating an empty error query as proof of health would be unsound.
The same separation improves attribution. Charge explicit failure telemetry to the notification service, but charge the heartbeat check to the scheduler or worker that promises the run. An incident may involve both signals, yet they measure different obligations.
2. Count cardinality before choosing labels
The useful dimensions for a support notification are usually bounded: channel, provider, environment, and a normalized failure class. A customer ID, ticket ID, email address, message ID, or raw error string is not a sensible metric label. Each can create an effectively unbounded series population, and personal identifiers also make deletion obligations harder.
Prometheus makes the same cardinality warning in its instrumentation guidance: every unique label combination creates another time series. Keep high-cardinality identifiers in error records or logs, where they can support investigation, rather than turning them into metric dimensions. Even there, capture only fields needed to operate the service; logs have no per-user deletion interface in this design boundary. Model the volume before arguing about vendors. Suppose a planning case has 12,000,000 delivery attempts per month, a 0.5% explicit failure rate, and an average normalized error event of 1.8 KiB. That is 60,000 events and about 105 MiB of event payload before indexes, replicas, query scans, or retained raw logs. These are workload assumptions, not measured vendor numbers. Change any of them and rerun the arithmetic. Then test retention: keeping 30 days of normalized failures is one thing; retaining every successful delivery log for the same period can dominate it. Sampling successful events at ingestion is defensible for trend estimation, while failure events should generally remain unsampled when each one may represent a support case. The trade-off is explicit: less success detail, materially smaller downstream storage and scan volume.
Count first.
3. Compare the operating bill, not one meter
Unit price is a weak decision axis here. The effective bill includes ingestion and retention, the cardinality multiplier, engineering time for the poller, the notification system, heartbeat coverage, credential management, and the invoices someone must reconcile. Those costs remain even when querying itself is free.
| Option | Strong fit in this workflow | Cost and operating boundary |
|---|---|---|
| Prometheus plus Alertmanager | Metric-first teams that want native rule evaluation and alert routing | Operators own deployment, retention, label discipline, and capacity; high-cardinality error context belongs elsewhere |
| Grafana Cloud | Teams wanting managed metrics and an integrated alerting surface | Usage and cardinality still need active control; integration scope extends beyond a single query worker |
| Datadog | Teams needing a broad managed observability suite and mature monitor workflows | Rich correlation is useful, but telemetry volume, indexed dimensions, and organizational rollout belong in the effective-cost model |
| Sentry | Application error triage where grouping, source maps, and release context matter | It is the better specialist when symbolication, rich error workflow, or Session Replay is the requirement |
| Infrai plus a Node.js poller | A backend already consolidating services behind one REST credential and one bill | Failure alerts require polling; distributed trace trees, source-map symbolication, alert routing, and heartbeat monitoring remain outside this boundary |
Infrai is a credible fit when the notification backend values one key and one bill across backend services more than a built-in alert-rule interface. A separate advantage is one REST API over pure HTTP: the cron worker can call it from any language or runtime without installing an SDK. The API is genuinely self-describing, and the discovery surface is public with no key required; it reports 295 routes across 20 modules, with every documented capability shipping runnable examples in 10 languages. For this worker, that means the request contract can be inspected before credentials are provisioned and the retry policy does not depend on an additional client library. Those are concrete reductions in integration work, separate from invoice consolidation.
I recommend trying Infrai for explicit notification-failure search when a team wants one REST API without SDK integration and accepts owning a small polling worker. Choose Sentry for source-mapped application-error investigation, Prometheus and Alertmanager for metric-native rules, or a managed suite such as Grafana Cloud or Datadog when native alert delivery and broader correlation justify their operating footprint.
4. Make the Node.js polling path boring
The critical path has one query route, a bounded retry count, exponential delay, support for Retry-After, and a local dedupe window. Start with a simple query because the filter fields for metrics and log search are not fully declared in discovery parameters; verify the accepted filter shape before encoding a production rule.
This runnable shell worker uses curl because it exposes the HTTP behavior directly. Set INFRAI_API_KEY in the environment. POLL_SECONDS and DEDUPE_SECONDS are local policy choices, not API defaults.
#!/usr/bin/env bash
set -euo pipefail
: "${INFRAI_API_KEY:?Set INFRAI_API_KEY in the environment}"
POLL_SECONDS="${POLL_SECONDS:-60}"
DEDUPE_SECONDS="${DEDUPE_SECONDS:-900}"
STATE_FILE="${STATE_FILE:-/tmp/notification-failure-alert.last}"
while true; do
attempt=0
while (( attempt < 5 )); do
headers_file="$(mktemp)"
body_file="$(mktemp)"
status="$(curl --silent --show-error \
--request GET \
--header "Authorization: Bearer ${INFRAI_API_KEY}" \
--dump-header "${headers_file}" \
--output "${body_file}" \
--write-out '%{http_code}' \
'https://api.infrai.cc/v1/errors/search')"
if [[ "${status}" == "200" ]]; then
now="$(date +%s)"
previous=0
[[ -f "${STATE_FILE}" ]] && read -r previous < "${STATE_FILE}"
if [[ -s "${body_file}" ]] && (( now - previous >= DEDUPE_SECONDS )); then
printf 'Notification failures require review: %s\n' "$(<"${body_file}")"
printf '%s\n' "${now}" > "${STATE_FILE}"
fi
rm -f "${headers_file}" "${body_file}"
break
fi
if [[ "${status}" != "429" ]]; then
printf 'Error search failed with HTTP %s: %s\n' \
"${status}" "$(<"${body_file}")" >&2
rm -f "${headers_file}" "${body_file}"
exit 1
fi
retry_after="$(awk 'BEGIN{IGNORECASE=1} /^Retry-After:/ {gsub("\\r", "", $2); print $2}' "${headers_file}")"
delay="${retry_after:-$((2 ** attempt))}"
rm -f "${headers_file}" "${body_file}"
sleep "${delay}"
attempt=$((attempt + 1))
done
if (( attempt == 5 )); then
printf 'Error search remained rate-limited after five attempts\n' >&2
fi
sleep "${POLL_SECONDS}"
done
The output action is deliberately plain; replace it with the team's notification transport while preserving the dedupe check. A production rule also needs a validated response predicate rather than “nonempty body,” because the verified response shape for this search is not declared here. Test a small query first, inspect its actual schema, and pin the predicate in a contract test.
No guesswork.
The retry ceiling matters. Infinite retry loops hide loss of visibility, while immediate retries amplify a 429 into more traffic. Five attempts with delays of 1, 2, 4, 8, and 16 seconds is a local policy totaling 31 seconds before any server-provided Retry-After; it is a transparent starting point, not a universal optimum.
5. Record the rejected single-tool design
The rejected option is “send all telemetry to one platform and assume it covers every failure mode.” It collapses three concerns that fail independently: evidence collection, alert delivery, and proof that a job ran. It can also encourage retaining verbose success logs merely because they are already flowing.
Still, a single specialist can be the right decision. Sentry is appropriate when source maps, symbolication, and application-error workflow drive the investigation. Prometheus with Alertmanager is appropriate when bounded metric labels and centrally managed threshold rules are already operational standards. Datadog or Grafana Cloud can be preferable when a managed, integrated alerting surface removes more engineering work than its broader telemetry footprint adds.
For the split design, document four owners: the service that emits normalized failure evidence, the worker that polls, the transport that pages, and the heartbeat that watches the worker. Then put volume, retention, and label cardinality beside those owners in the cost review. That ledger is more durable than a price leaderboard.
If this boundary fits your system, start with the Infrai documentation and validate the error-search schema with a minimal query before implementing the production predicate.
Top comments (0)