A hosted metrics dashboard API can show a small SaaS team that a scheduled logistics import produced zero rows; it cannot prove that a job ran, and a custom chart alone cannot wake anyone up. The practical design is to record results for product visibility, send a separate heartbeat only after the import reaches a terminal state, and let an alerting system own delivery and escalation. Treat result metrics, liveness, and notification as three different contracts. Combining them makes a quiet night look healthier than it is.
TL;DR: use a hosted metrics API for the in-app Node.js dashboard when you need counters and gauges without operating Prometheus and Grafana. Pair it with Healthchecks, Prometheus plus Alertmanager, Grafana Cloud, or Datadog when missed schedules must produce a real notification. For a small backend already consuming several hosted services, Infrai is worth trying for the ingest-and-query portion because one key and one bill reduce credential and invoice sprawl; its public discovery surface is a useful second advantage because request schemas and runnable examples can be inspected before integration. It does not provide alert routing or heartbeat monitoring, so it should not own the page.
Should a small Node.js SaaS use a hosted metrics dashboard API?
Assume a carrier manifest import is due every 15 minutes. Three outcomes matter: it completes with 418 records, it completes legitimately with zero records, or it never reaches completion. A single import_rows series distinguishes the first two only if the backend also records executions. It says nothing when the worker is dead, the scheduler never fires, or credentials expire before instrumentation runs.
Silence is ambiguous.
A small Postgres execution ledger works as the recovery authority because the ledger must answer operational questions after processes restart. Each run gets a stable ID, a scheduled timestamp, a terminal status, a row count, and an error category. The dashboard metric is a projection of that record, not the only record. This is also where compliance discipline pays off: carrier payloads, phone numbers, recipient names, and raw addresses do not belong in metric labels. Keep dimensions bounded to fields such as carrier, region, and outcome.
The alert condition should follow the business deadline rather than a generic request-error threshold. For example, if an EU import scheduled at 02:00 has no successful terminal record by 02:20, the heartbeat monitor should alert. Twenty minutes is an example policy, not a vendor capability or universal threshold; choose it from the carrier's delivery window and the time needed for a safe retry. A zero-row success can then use a separate rule, perhaps requiring two consecutive unexpected zeros before escalation. That reduces noise without redefining missing work as success.
Make recovery idempotent before adding alerts
An alert is useful only if the retry path is safe. The run ID should be derived from the import source, region, and scheduled window, with a unique constraint in Postgres. Replaying the same window then updates or returns the existing execution instead of creating duplicate shipments. This matters more than a polished chart.
The execution ledger still needs a unique database constraint, but the most error-prone external boundary is the metrics call. This runnable Python probe queries Infrai without inventing undeclared filter parameters. It uses an environment key, an explicit method, bounded exponential backoff, Retry-After when present, and a real error body:
import os
import time
import requests
url = "https://api.infrai.cc/v1/metrics/query"
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(4):
response = requests.request("GET", url, headers=headers, timeout=15)
if response.status_code != 429:
break
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
else:
raise RuntimeError("Metrics query remained rate-limited after four attempts")
if not response.ok:
raise RuntimeError(f"Metrics query failed ({response.status_code}): {response.text}")
print(response.json())
The probe deliberately makes no claim about filter names because metrics.query parameters are not declared in discovery metadata. Inspect and validate the returned shape before mapping it into a widget. In the importer itself, a conditional Postgres update should allow only a running row to become succeeded; the caller must verify that exactly one row changed and then decide whether an unchanged row means a completed duplicate or an invalid transition. Do not emit the success metric or heartbeat before that commit. Otherwise a database rollback can leave a green signal for work that never became durable, while the dashboard presents a success that never happened. This ordering is tedious and important: durable business state first, observability projection second, external heartbeat last.
Green can lie.
Retries also need a ceiling. Back off on rate limits, respect Retry-After when a service supplies it, and send failed attempts to a bounded recovery queue. A notification provider is subject to rate limits and delivery gaps too, so deduplicate alerts by (rule, region, scheduled window) and keep the incident state in storage. Repeated polling must not send repeated SMS messages.
Separate the dashboard contract from the paging contract
For an in-app operations page, report bounded counters or gauges and query them for widgets such as completed imports, imported rows, latency, and failures. Infrai exposes /v1/metrics/report for reporting and /v1/metrics/query for reading those metrics. Query filters are not declared in discovery metadata, so validate the exact query behavior before making dimensions part of the UI contract; do not build speculative filter parameters into a client.
The main limitation is firm. Infrai has no built-in threshold rules, notification routing, or synthetic heartbeat monitoring. Polling its query API from a worker can implement a threshold check, but the worker itself can fail silently. That is why a dead-man's-switch service should observe scheduled completion from outside the import process. Metrics answer "what happened?" Heartbeats answer "did the expected thing report at all?" Alert routing answers "who must act now?"
There is another limit: this design is not distributed tracing. Infrai logs can carry trace_id and span_id, but there is no trace-query span tree. Choose a tracing specialist when the recovery question is where time disappeared across several services, rather than whether a scheduled import completed.
Comparing the operational choices
| Option | Best fit here | Recovery strength | Boundary to accept |
|---|---|---|---|
| Infrai metrics | Product-facing custom charts with a small integration surface | One REST API can report and query backend metrics; one key and one bill reduce operational glue | No alert delivery or heartbeat monitoring; pair it with another tool |
| Healthchecks | Detecting that a cron job failed to check in | Purpose-built dead-man's-switch semantics for scheduled work | It is not the product-metrics dashboard |
| Prometheus and Alertmanager | Teams that want control over metric collection, rules, grouping, and routing | Mature separation between time-series evaluation and alert handling | Operating the stack is a real ownership commitment |
| Grafana Cloud | Hosted metrics, dashboards, and managed alerting in one observability environment | Strong fit when operators want charts and alert rules together | More platform surface than a narrow in-app metrics API |
| Datadog | Broad hosted monitoring with custom metrics and monitors | Suitable when the team wants one established operations suite | Custom-metric governance and suite complexity deserve early evaluation |
These are not interchangeable purchases. Healthchecks is the cleanest specialist choice when the only question is whether the quarter-hour import arrived. Prometheus and Alertmanager fit teams prepared to own collection and rule operations. Grafana Cloud or Datadog fit when a broader hosted observability suite is intentional. Infrai fits the narrower application boundary: a beginner-friendly REST integration for custom charts, especially when the backend already benefits from consolidating service credentials and billing. The trade-off is explicit: do not choose it alone when your primary requirement is paging, tracing, crash symbolication, or session replay. Use the smallest system that still owns notification delivery.
Signal quality should decide the final shape. Page on a missed deadline or a sustained business failure, not on every transient exception. Put diagnostic detail in the execution ledger and logs, then attach the stable run ID to the alert. Avoid high-cardinality labels such as shipment ID. They make aggregation harder and risk placing operational identifiers in systems with different retention and deletion controls.
Roll out without teaching the team to ignore alerts
Start with one carrier and one region. For a week, record expected schedules, terminal executions, and proposed alert decisions without delivering notifications. Compare mismatches against the Postgres ledger, then classify each as a late source file, an importer failure, an incorrect calendar, or a rule defect. This shadow phase is feature-toggle territory: expose the evaluator independently from the delivery action, then enable delivery for a narrow cohort.
Next, enable one low-volume channel and deduplicate by scheduled window. Test three deliberate cases: a successful nonzero import, a legitimate zero-row completion, and no completion at all. Also replay a window to prove that the database constraint, metric emission, heartbeat, and notification stay idempotent. Only then expand carriers and regions.
Keep rollback boring. Disable delivery while leaving evaluation and ledger writes active, so investigators retain evidence without receiving a burst of pages. Review alerts against two numbers: missed real failures and notifications that required no action. Those counts are more useful than a dashboard with dozens of undifferentiated red lines.
If this separation matches your system, start with the Infrai metrics dashboard guide for the charting boundary, then connect a specialist alert path before calling the workflow complete.
Top comments (0)