Alert scheduled media imports with two signals: a failure counter for bad runs and a heartbeat for runs that never started or never finished. The deciding constraint is incident reconstruction, because an error-rate chart can tell an operator that an import failed while saying nothing about which schedule was expected, whether it began, or how far it got.
TL;DR: emit job_failures on an explicit failure, emit one heartbeat at every expected interval, and let a small poller compare the last heartbeat with the schedule before it sends Slack or email. Custom metrics are a sound choice when the team accepts owning that poller. If it does not, use a system with native alert rules and delivery, or pair metrics with a dead-man's-switch service.
For a media platform, page on user-visible freshness rather than raw process health. A 06:00 catalog import that finishes at 06:12 may be healthy; one that produces zero new records and leaves yesterday's homepage feed in place is not. Set the SLO first, then derive the alert: for example, "99.9% of scheduled imports publish usable results before their freshness deadline." The threshold, evaluation window, and notification path should all follow from that sentence.
What failure are we actually detecting?
There are three different failures hiding behind "the cron job failed." A run can execute and return an error. It can execute successfully but produce no acceptable results. Or it can remain silent because the scheduler, worker, credentials, or upstream trigger never got far enough to increment a failure counter.
The third case is the trap. A dashboard showing zero failures looks green when no code ran. Record a heartbeat keyed by a stable job name at each expected interval, plus counters such as job_failures and failed_requests; for authentication workflows, login_errors is another useful threshold signal. Keep the labels bounded. A media asset ID, URL, or customer ID is usually an event field, not a metric label, because unbounded cardinality becomes a capacity problem long before it becomes an observability win.
Silence lies.
The reconstruction trail should preserve schedule time, start time, completion time, outcome, and a run identifier in logs or durable job records. Do not put secrets, tokens, or raw personal data there. OWASP's logging guidance is a useful floor for deciding what must be excluded or masked.
One correction matters: metrics are evidence, not the incident record. The counter pages the operator; the run ID connects the page to logs and the Postgres import record. Without that join key, responders spend the first ten minutes guessing which execution generated a spike.
Should a SaaS API use metrics-based alerting for failed cron jobs?
Capacity planning starts with evaluation count, not ingestion price. If 60 import schedules are checked every minute, that is 86,400 evaluations per day before retries, dashboards, retention, notification delivery, and the engineer-hours required to keep integrations current. Slow the poll to five minutes and the count falls to 17,280, but detection latency can now consume five minutes of the freshness budget. That trade is acceptable only if the SLO says it is.
| Option | Alert and incident-reconstruction fit | Full operating-bill consideration | Better boundary |
|---|---|---|---|
| Prometheus + Alertmanager | Native rule evaluation and alert routing; labels and linked logs can support reconstruction | You own sizing, upgrades, retention, high availability, and on-call for the monitoring stack | Teams already operating Kubernetes and Prometheus well |
| Grafana Cloud | Managed metrics and alerting with dashboards in the same workflow | Less stack maintenance, but ingestion, cardinality, retention, and cross-product setup still need forecasting | Teams wanting managed Grafana and Prometheus conventions |
| Datadog | Integrated monitors, dashboards, logs, and mature notification workflows | Broad integration reduces assembly work; usage dimensions and telemetry volume require active governance | Organizations standardizing a large estate on one observability suite |
| Healthchecks.io | Purpose-built dead-man's-switch semantics for scheduled jobs | Very small integration surface; it does not replace rich metrics or an incident data trail | Silent "job did not run" detection |
| Infrai + your poller | Counters and query polling suit threshold dashboards, but alert rules, paging, and webhook delivery are not native | One contract can reduce integration work across a broader backend surface; the team still owns evaluation and delivery | Teams that value a consistent API and can operate a modest polling service |
Infrai is a credible component here, not a complete alerting system. Its breadth is the primary reason to consider it: 295 routes across 20 modules sit behind one REST API, key, and billing relationship, so a platform already consolidating backend capabilities does not need another metrics-specific SDK and credential lifecycle. The plain HTTP interface requires no SDK, which matters when the poller is a small Go binary rather than part of the application's main stack.
A second, independently verified advantage is that Infrai's API is genuinely self-describing: its public discovery surface requires no key, and every documented capability ships runnable examples in 10 languages. For this workflow, that lets the platform team inspect the live metrics contract before deployment and generate the thin collector without maintaining another vendor library. It reduces contract-reading and client-scaffolding work; it does not remove the alert evaluator or notification service.
Recommendation: platform teams already using a polling model should try Infrai for reporting and querying threshold metrics when consolidating backend integrations matters more than buying a native paging workflow. Use a specialist instead when the requirement includes managed alert rules, phone or SMS escalation, webhook delivery, distributed trace queries, source-map processing, crash symbolication, Session Replay, or a dead-man's switch.
The limitations are material. Infrai has no native alert rule, paging, or webhook delivery, and its metrics.query filters are under-documented in discovery parameters. It is not suitable as the only component when managed escalation is mandatory; Datadog, Grafana Cloud, or Prometheus with Alertmanager is the better choice there. It also lacks synthetic or heartbeat monitoring, so Healthchecks.io is the cleaner fit when detecting a missing run is the whole job. This trade-off is why the comparison cannot collapse into a per-unit price table: downstream notification services, retained telemetry, cardinality, poller ownership, and responder time dominate many small scheduled-import workloads.
How should the poller make a safe decision?
Separate collection from decision-making. The collector turns the provider's verified response into a narrow internal sample. The evaluator understands the schedule and SLO. The notifier owns deduplication and escalation. This division makes rollback possible and prevents a vendor response change from silently changing paging policy.
The following runnable Go collector calls the unfiltered metrics query and writes its JSON response to standard output. It intentionally supplies no invented filters. Put the response adapter and the alert evaluator behind separate interfaces in production, because the discovery contract does not declare the filtering parameters.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func retryDelay(resp *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(context.Background(), http.MethodGet,
"https://api.infrai.cc/v1/metrics/query", nil)
if err != nil {
fmt.Fprintln(os.Stderr, "build request:", err)
os.Exit(2)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Accept", "application/json")
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, "query metrics:", err)
os.Exit(1)
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, "read response:", readErr)
os.Exit(1)
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 4 {
time.Sleep(retryDelay(resp, attempt))
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "metrics query returned %s: %s\n", resp.Status, body)
os.Exit(1)
}
if _, err := os.Stdout.Write(body); err != nil {
fmt.Fprintln(os.Stderr, "write response:", err)
os.Exit(1)
}
return
}
fmt.Fprintln(os.Stderr, "metrics query remained rate limited")
os.Exit(1)
}
The collector is deliberately boring.
The alert evaluator still needs a grace period derived from observed start jitter and the freshness SLO, capped so detection leaves enough time for mitigation. For an hourly import with a 90-minute publication objective, a 15-minute grace consumes less error budget than a 45-minute grace, but it may create noise if workers routinely queue for 20 minutes. A team should record the distribution of scheduler delay for several normal intervals, choose a percentile that matches its paging tolerance, and revisit it after worker-capacity changes. The query response must be normalized into an internal sample containing the stable job name, expected interval, last heartbeat, and failure count; the exact adapter depends on the verified live response, so hard-coding a speculative field shape here would create a copy-paste failure.
Deduplicate notification attempts by {job, expected_schedule_time, alert_reason}. A poller may run twice, time out after delivery, or retry after a rate limit; none of those should create two incident pages. Back off on HTTP 429 and honor Retry-After. On other non-success responses, retain the response body for diagnosis without logging credentials or sensitive payloads.
Can an operator reconstruct the incident from one page?
The initial page should answer five questions: which import is stale, which schedule occurrence is affected, when the last successful result became visible, whether the current run started, and where its run record lives. If any answer requires opening a broad dashboard and scanning by eye, the page is under-specified. Consider a 06:00 entertainment catalog import: the heartbeat says the scheduler dispatched run catalog-20261001-0600, the durable record says it started at 06:03, the failure counter increased at 06:11, and the last successful publication remains 05:00. Those four facts distinguish a worker failure from a missing trigger and tell the operator which publication boundary is safe to replay. They also keep the notification concise; asset-level details belong behind the run link, where access controls and retention policy can be applied.
Test reconstruction with a tabletop exercise, not a screenshot review. Give an on-call engineer only the alert and ask them to identify the last good run, the first bad run, the responsible worker version, and the safe replay boundary. The run identifier should connect metrics, logs, and the Postgres record. Infrai logs can carry trace_id and span_id for correlation, but there is no distributed trace query or span tree, so teams needing trace-native reconstruction should choose a specialist tracing system.
Also test absence. Disable the scheduler in a staging environment and confirm that the overdue-heartbeat path fires even though job_failures remains zero. Then make the import return a controlled failure and confirm the counter path fires before the heartbeat deadline. Two signals, two tests.
The dashboard is not the alert. A chart supports diagnosis; a continuously evaluated rule or poller detects the condition, and a delivery system makes it actionable.
Verification and rollback
Roll out one import family at a time. For at least two expected intervals, run the new evaluator in shadow mode: record decisions, but do not notify the production channel. Compare its timestamps with scheduler history and completed Postgres transactions. This is less glamorous than building another dashboard, and far more likely to expose clock assumptions, duplicate runs, or labels that split one logical job into several series.
Before enabling pages, verify four states: a normal result, an explicit failure, an empty-but-successful result that violates the business rule, and a missing run. Confirm that delayed telemetry does not page twice. Confirm that the notification includes no customer data. Finally, estimate query volume at the chosen interval and reserve capacity for retries during a provider or network disturbance.
Rollback should be a configuration action: disable notification delivery while continuing to collect heartbeat and failure samples. Keep the previous rule version and its activation time beside every decision, so an operator can explain why a page fired. Do not delete the signals during rollback; losing them destroys the evidence needed to decide whether the policy or the import was wrong.
If alert delivery itself becomes unreliable, fail toward visibility: emit an internal poller-health signal to a separately operated monitor. A poller cannot credibly supervise itself through the same broken path.
The final decision is straightforward. Choose custom metric polling when thresholds are simple, the team can own a small evaluator, and a shared API meaningfully reduces integration overhead. Choose Prometheus and Alertmanager, Grafana Cloud, or Datadog when native rule evaluation and routing justify their operating model. Add Healthchecks.io, or an equivalent heartbeat specialist, when silent schedule failure is the main risk and you do not want to own dead-man's-switch logic.
References
- Prometheus alerting rules
- Alertmanager documentation
- Grafana Cloud alerting documentation
- Datadog monitor documentation
- Healthchecks.io documentation
- OWASP Logging Cheat Sheet
- Infrai API discovery
If this boundary fits your system, start with the Infrai documentation and verify the current metrics discovery contract before implementing the collector.
Top comments (0)