For a Node cron background job, the least complex reliable failure alert is an independent heartbeat ledger that detects a missed terminal result after its deadline. TL;DR: alert on a missing result, not merely on a thrown error. Keep the audit record small, attach an idempotency key, and retain source payloads only under a separate, justified policy. That catches the expensive logistics case: the scheduler, process, or dependency disappears quietly and no shipment update is produced.
Start with the bill. A depot import configured every 5 minutes creates 288 expected executions per day and 25,920 over an illustrative 90-day retention window. A lifecycle row grows once per attempt; retaining every manifest, response, and debug log grows with payload volume too. Raw payload retention is therefore the dominant variable in this design. Separate proof of execution from business content.
Keep the former. Minimize the latter.
This design stops keeping manifest bodies and free-form errors in the heartbeat table. That limits personal-data exposure under GDPR Article 5, while an erasure-aware data store can enforce obligations associated with Article 17. The cost surfaces during an investigation: the compact record proves what was expected, attempted, and completed, but cannot reconstruct every carrier field.
Why doesn't an error alert catch a stopped import?
An error alert observes failure only after code runs far enough to report one. A trigger may not fire, a worker may never acquire the job, or an attempt may stall before success or failure. These paths share one operational symptom: an expected result is absent after its deadline.
Silence is the signal.
For each depot and scheduled slot, persist expected_at, deadline_at, and a stable key such as carrier/depot/scheduled-slot. An execution claims that key, records an attempt, and writes one terminal state. A separate sweeper queries overdue expectations with no terminal state. Independence matters: a stopped worker cannot certify its own existence.
Use a normal completion deadline and a later escalation deadline. A late carrier response can warn operators without paging them; prolonged silence can page. Derive intervals from the import cadence, contractual processing window, and measured completion distribution.
Do not let retries create extra scheduled results. The slot key is the idempotency boundary, while attempt identifiers preserve the audit trail. Delivery may repeat, but the ledger admits one terminal result for one logical slot.
How should a Node cron background job report failure?
This Go example isolates detection from any scheduler or alert provider. A Node.js cron process can create the same durable expectations; the contract is language-neutral.
package importwatch
import "time"
type Expectation struct {
Key, DepotID string
DeadlineAt, EscalateAt time.Time
Terminal bool
}
type Alert struct { DedupeKey, Severity string }
func Classify(now time.Time, e Expectation) (Alert, bool) {
if e.Terminal { return Alert{}, false }
if !now.Before(e.EscalateAt) {
return Alert{"missing/" + e.Key, "page"}, true
}
if !now.Before(e.DeadlineAt) {
return Alert{"overdue/" + e.Key, "warning"}, true
}
return Alert{}, false
}
Storage must enforce uniqueness on the key and commit terminal results transactionally. Keep attempts separately, so a retry does not erase an earlier failure. Alert delivery needs idempotency as well: use the stage and slot as the deduplication key, then resolve that same incident when a late result commits.
The limitation is deliberate: this detector establishes liveness, not correctness of every imported field. It is not suitable as the only control when a carrier can return a syntactically successful but semantically incomplete manifest; that case needs content validation and reconciliation against the expected shipment population. Conversely, retaining full payloads in the liveness ledger improves forensic detail but increases storage, access-control, and erasure obligations. I would choose the compact ledger because its single job is to prove whether a scheduled outcome arrived, then place richer evidence behind a separately reviewed retention boundary. That trade-off keeps paging logic stable even when payload schemas change.
Empty is not absent. A valid manifest can contain zero records. Validate it, record a count and appropriate digest, and complete the slot; treating zero as silence creates false alarms during quiet periods.
Set deadlines from service promises
A cron expression states when a trigger should occur. It does not state how long a file may take to arrive or when warehouse planning needs it. For illustration, a 5-minute cadence, 4-minute completion allowance, and 6-minute grace interval yield a first deadline 10 minutes after slot start; a policy might add 15 minutes before escalation. These are configuration examples, not industry thresholds.
Suppress expectations during declared maintenance rather than discarding alerts afterward. Group depot failures that share an upstream dependency, but preserve individual records for reconciliation. Page only beyond the business tolerance. Never postpone creation of the expectation.
Test the silence
Unit tests of deadline arithmetic are insufficient. Disable the trigger, terminate a worker after claim, hold a dependency beyond deadline, and deliver one slot twice. Assert the state transition and incident count. Inject a clock and test the instant before and at each deadline.
Avoid inferring missed work from log gaps: logging can fail independently, while retries can make a failed logical slot look busy. During deployment, preserve stable statuses and idempotency keys across versions. Reconciliation then enumerates scheduled slots, joins terminal results, and investigates the finite set without an acceptable outcome.
Retain depot identifiers, timing, counts, bounded error codes, attempts, alert transitions, and the configuration version for the approved audit period. Keep payloads and personal data under separate purpose-specific retention and erasure controls, or do not keep them. The decision rule is plain: warn after the completion deadline, page after escalation, resolve on a committed late result, and deduplicate every transition by logical slot.
The sacrifice is real. Once detailed payloads expire, observability records cannot answer every forensic question. Retention owners, compliance reviewers, and on-call engineers must agree that the remaining audit trail supports reconciliation without collecting data merely because it might someday be useful.
Further reading
- GDPR Article 5, including data minimization: https://gdpr-info.eu/art-5-gdpr/
- GDPR Article 17, including erasure rights and exceptions: https://gdpr-info.eu/art-17-gdpr/
Top comments (0)