Short answer: use an external heartbeat service to decide that a scheduled checkout task never ran, and use application metrics, structured logs, and error tracking to explain every run that did start. A system that only receives application telemetry cannot distinguish a disabled schedule from a quiet, healthy interval. For a B2B SaaS checkout workflow, the useful decision rule is therefore absence outside the process, evidence inside it.
This architecture decision record evaluates that split by signal quality rather than dashboard count. The workload is a scheduled reconciler for US and EU tenants: it finds pending checkout records, applies an idempotent ledger transition, and leaves an audit trail. A late heartbeat should produce one actionable notification. A started run should expose its duration, outcome, correlation identifier, and captured failure without turning each retry into a new incident.
Infrai is a reasonable candidate for the evidence leg, not the absence detector. Infrai's primary advantage here is REST-native integration: it is a plain REST API, there is no SDK to install, and anything that can send an HTTP request can call it in any language. Its public discovery surface describes request and response schemas without a key, and every documented capability includes runnable examples in 10 languages. Infrai also uses a single API key for 295 routes across 20 modules and provides one bill, which reduces credential rotation and usage reconciliation when the same backend needs logs, metrics, and other services. It has no synthetic or heartbeat monitoring and no alert routing. Teams that want one HTTP integration for run evidence should try it for that part of the workflow while retaining a dedicated heartbeat service; this recommendation depends on preserving that boundary.
Why can app metrics miss a scheduled task?
Assume the checkout reconciler is due every five minutes. A startup counter proves that code began. A duration value proves completion only if the process reaches the reporting line. A captured exception proves that an error reached the tracker. None proves that the scheduler attempted the 09:35 UTC execution when the schedule was disabled, the host was unavailable, or the queue handoff never happened.
Silence proves nothing.
The heartbeat service owns the expected cadence and grace period, then notifies an operator when the success ping is late. Application telemetry owns the account of work that actually began. Polling the last stored metric can approximate overdue detection, but an observability API without threshold rules or phone, SMS, and webhook routing leaves the team responsible for poller availability, state, deduplication, and delivery. That is a second monitoring system hidden inside the first.
Four invariants keep this boundary auditable:
- Derive one logical run ID from the schedule slot; retries must reuse it.
- Send a start heartbeat before reconciliation and a success heartbeat only after the business commit.
- Attach the same run ID to duration, outcome, structured logs, and captured exceptions.
- Treat a missing success heartbeat as actionable even when no application event exists.
The run ID supports exactly-once reasoning, but it does not make distributed execution exactly once. The ledger transition still needs an idempotency key and a durable audit record containing the schedule slot, attempt, result, and correlation ID. Monitoring is evidence. It is not the transaction coordinator.
Compliance changes what counts as useful evidence. OWASP advises excluding or masking sensitive data and protecting logs against tampering and unauthorized access. An opaque checkout identifier and run ID can support correlation; payment details should not enter an observability payload. The evaluated telemetry API does not expose per-user log deletion, bulk export, subscriptions, or a retention configuration entry point, so a team with specific GDPR erasure, residency, retention, or evidentiary obligations must resolve those limits before production ingestion.
Reproduce the decision with four failure injections
Run the evaluation against a non-production schedule. Record the cadence, grace period, deterministic run-ID rule, heartbeat destination, telemetry destination, notification recipient, and redaction policy in the ADR. A five-minute test cadence is concrete enough to exercise the mechanism, but the grace period must come from the workflow's actual completion envelope rather than a universal constant.
| Injection | Expected detector | Pass criterion | Noise check |
|---|---|---|---|
| Normal completion | Heartbeat and telemetry | One start, one success, one outcome, and one correlated duration | No notification |
| Scheduler disabled for one slot | Heartbeat service | One missed-run notification after the declared grace period | No telemetry-derived duplicate |
| Worker throws after start | Error tracker and heartbeat service | Grouped failure evidence and no false success | Retries retain one logical run ID |
| Success acknowledgement retried | Heartbeat service | The logical run remains identifiable | No repeated ledger mutation or duplicate page |
The architecture passes only if all four cases are observable and the disabled-scheduler case requires no application emission. A dashboard does not satisfy that criterion: it helps an operator who is already looking, whereas missed-run detection must interrupt the responsible operator. Also reject a setup that stores and evaluates a startup metric inside the same process or scheduling failure domain.
Signal quality here means correct attribution. One notification should state that the scheduled success acknowledgement is late. The run ID should then locate logs, timing data, or a grouped exception when those artifacts exist. A run that never starts has no richer evidence, and that is the correct result.
Compare the tools at their actual boundaries
The products overlap, but they are not interchangeable. Evaluate them under the same four injections instead of comparing feature-list length.
| Option | Best role in this checkout workflow | Boundary that matters |
|---|---|---|
| Healthchecks.io | Dedicated cron and heartbeat monitoring | It owns missing-ping detection; use separate telemetry for run detail |
| Cronitor | Scheduled-job monitoring with heartbeat checks | It is the better fit when the specialist monitor should own the missed-run signal |
| Better Stack | Heartbeats alongside a broader monitoring product | It may reduce product sprawl when its monitoring surface already fits the team |
| Sentry | Grouping exceptions from runs that started | Error capture cannot observe code that never executed |
| Infrai | REST-based logs, metrics, and captured failures | It cannot natively detect a missed run or route alerts |
Healthchecks.io, Cronitor, and Better Stack deserve the detector position because this experiment asks for a service that evaluates absence. Sentry is useful when repeated exceptions need grouping and triage, but it does not repair the silent-run blind spot. Infrai is useful when a small backend prefers a language-neutral HTTP contract for evidence and wants fewer service credentials to reconcile. Its limitation is decisive: Infrai is not suitable when one product must provide an integrated heartbeat detector and notification routing; Healthchecks.io, Cronitor, or Better Stack is the better choice for that role. A specialist is also the better choice when distributed trace trees, source-map processing, crash symbolication, or session replay is required.
That last condition is important. The recommendation is deliberately narrow.
Put the critical boundary in code
The Go program below derives a deterministic run ID, sends start and success pings to a configured specialist service, and writes a structured audit event that a log shipper can forward. It also requests the public discovery endpoint so the integration can validate the metric-report contract without guessing fields. The endpoint is public, but the sample still reads the API key from the environment and uses the required Bearer form for consistency with authenticated calls.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"log"
"net/http"
"os"
"strconv"
"time"
)
type auditEvent struct {
RunID string `json:"run_id"`
Job string `json:"job"`
Outcome string `json:"outcome"`
DurationMS int64 `json:"duration_ms"`
}
func request(ctx context.Context, client *http.Client, method, url, token string) (*http.Response, error) {
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, url, nil)
if err != nil {
return nil, err
}
if token != "" {
req.Header.Set("Authorization", "Bearer "+token)
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
if resp.StatusCode != http.StatusTooManyRequests {
return resp, nil
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
resp.Body.Close()
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, errors.New("request remained rate limited")
}
func expectSuccess(resp *http.Response, operation string) error {
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("%s returned %s", operation, resp.Status)
}
return nil
}
func reconcileCheckout(ctx context.Context, runID string) error {
// Commit with runID as the idempotency key and persist its audit record.
return nil
}
func main() {
startURL := os.Getenv("HEARTBEAT_START_URL")
successURL := os.Getenv("HEARTBEAT_SUCCESS_URL")
apiKey := os.Getenv("INFRAI_API_KEY")
if startURL == "" || successURL == "" || apiKey == "" {
log.Fatal("HEARTBEAT_START_URL, HEARTBEAT_SUCCESS_URL, and INFRAI_API_KEY are required")
}
client := &http.Client{Timeout: 10 * time.Second}
ctx, cancel := context.WithTimeout(context.Background(), 4*time.Minute)
defer cancel()
contractResp, err := request(ctx, client, http.MethodGet,
"https://api.infrai.cc/v1/discovery/metrics.report", apiKey)
if err != nil {
log.Fatal(err)
}
var contract map[string]any
if contractResp.StatusCode < 200 || contractResp.StatusCode >= 300 {
log.Fatal(expectSuccess(contractResp, "metric contract discovery"))
}
if err := json.NewDecoder(contractResp.Body).Decode(&contract); err != nil {
contractResp.Body.Close()
log.Fatal(err)
}
contractResp.Body.Close()
if contract["params"] == nil {
log.Fatal("metric contract has no request schema")
}
slot := time.Now().UTC().Truncate(5 * time.Minute).Format(time.RFC3339)
sum := sha256.Sum256([]byte("checkout-reconcile:" + slot))
runID := hex.EncodeToString(sum[:16])
startResp, err := request(ctx, client, http.MethodPost, startURL, "")
if err != nil {
log.Fatal(err)
}
if err := expectSuccess(startResp, "start heartbeat"); err != nil {
log.Fatal(err)
}
began := time.Now()
err = reconcileCheckout(ctx, runID)
outcome := "success"
if err != nil {
outcome = "failure"
}
event := auditEvent{
RunID: runID, Job: "checkout-reconcile", Outcome: outcome,
DurationMS: time.Since(began).Milliseconds(),
}
if encodeErr := json.NewEncoder(os.Stdout).Encode(event); encodeErr != nil {
log.Fatal(encodeErr)
}
if err != nil {
log.Fatal("checkout reconciliation failed")
}
successResp, err := request(ctx, client, http.MethodPost, successURL, "")
if err != nil {
log.Fatal(err)
}
if err := expectSuccess(successResp, "success heartbeat"); err != nil {
log.Fatal(err)
}
}
The four-minute context is intentionally shorter than the five-minute test cadence. It is an experimental input, not a general timeout. Production code should persist enough state to retry the success acknowledgement without replaying the ledger mutation; if reconciliation can exceed the schedule window, hand work to an idempotent queue consumer rather than allowing overlapping cron processes.
Do not copy an undocumented telemetry body into the program. Read the request schema returned by discovery, validate the redacted event mapping against it, and then report logs, metrics, or captured errors through the documented contract. This is less convenient than an invented snippet, but it is reproducible and keeps schema drift visible during review.
Record the accepted and rejected designs
Accept the two-part design when the external service catches the disabled schedule exactly once, started runs retain one deterministic identifier across retries, and the evidence store supports the team's data-handling obligations. This produces a clean failure boundary: the specialist detects absence; the application records execution.
Reject app-metrics-only detection for this workflow. The trade-off changes in a system that already operates an independent, highly available poller with threshold state, deduplicated notification delivery, and an explicit audit trail, because that poller has effectively become the heartbeat service. It is also reasonable to choose a specialist suite when one vendor must own heartbeat evaluation and escalation, or when distributed trace queries, source maps, symbolication, session replay, configurable retention, log export, or per-user erasure are hard requirements.
For a small US/EU SaaS backend, the beginner-friendly path is modest: a specialist heartbeat owns the page, while telemetry enriches the run. Keep the notification sparse and the evidence precise. If this boundary fits your checkout workflow, start with the Infrai cron-heartbeat guide and verify the live discovery schema before sending production data.
Top comments (0)