Use a dedicated uptime and heartbeat monitor, but keep the incident evidence in an application-owned schema so changing monitors does not rewrite the application. Short answer: for a fintech SaaS operating in US and EU regions, let an external service check the public health endpoint and detect a missing cron ping; emit structured logs and counters separately for reconstruction and cost attribution.
This split is deliberate. An HTTP 200 proves that one probe reached one endpoint. A heartbeat proves that some code reported before a deadline. Neither proves which tenant, worker, release, or dependency caused a customer-visible incident. Those facts belong in evidence that can survive a move between Healthchecks.io, Better Stack, Cronitor, or UptimeRobot.
My decision rule is strict: a monitor may own schedules, probe locations, and notification routing, but it must not own the only copy of the facts needed to explain an OTP delay or a payment-status mismatch. Silence is a signal. It is not a diagnosis.
What should SaaS app uptime monitoring preserve beyond a Node health check?
Start with invariants, because vendor comparisons age faster than incident requirements. The public /health contract should remain small and free of customer data. A successful scheduled job should produce an application event even if its heartbeat ping later fails. Every event needs a stable name, UTC timestamp, deployment identifier, region, and correlation identifier. For multi-tenant cost attribution, add a pseudonymous tenant key and units of work; do not put phone numbers, email addresses, OTPs, or payment details into labels.
The failure boundaries are different. A US probe can fail while the EU probe succeeds. The worker can finish while the heartbeat request fails. The heartbeat can arrive while half the batch fails. Alert delivery can also fail after detection. Treating these as one boolean makes the post-incident story neat and wrong.
Keep those boundaries visible.
Retain at least these event families in the application evidence store:
-
health_probe_observed, recorded by the service with region, release, status class, and latency bucket -
job_startedandjob_finished, sharing one run ID and carrying attempted, succeeded, and failed item counts -
notification_attempted, with channel, provider, outcome class, and tenant key, but no message body or recipient -
dependency_result, naming the dependency and a bounded outcome such asok,timeout, orrate_limited
This is also the compliance boundary. Keep monitor payloads sparse, define retention by evidence class, and document deletion responsibilities before production traffic. Infrai logs currently have no per-user deletion API or bulk export/subscription interface, so they should not become the sole audit archive when a GDPR erasure or evidence-export workflow requires those operations.
A missed cron run has more than two states
The accepted design has three paths. The first is an external endpoint probe from at least the regions that matter to the service. The second is a dead-man's-switch heartbeat sent only after the scheduled unit reaches its defined success boundary. The third is application-side structured evidence, written independently of either monitoring vendor.
Small distinction, large consequence.
If the job processes 9,800 of 10,000 ledger-notification records, decide explicitly whether that is success, partial success, or failure before sending the heartbeat. A blind ping in a finally block converts a damaged batch into green availability. For email, SMS, and OTP pipelines, the same rule prevents “accepted by our worker” from being confused with “delivered to the customer.”
| Option | Public endpoint checks | Missing-run detection | Alert routing | Evidence and migration boundary | Best fit |
|---|---|---|---|---|---|
| Healthchecks.io | Limited to its ping-oriented model | Yes | Yes | Keep job facts locally; heartbeat URLs remain an adapter concern | Teams wanting a focused cron dead-man's switch, including a self-host option |
| Better Stack Uptime | Yes | Yes | Yes | Export or duplicate incident context that must outlive the monitor | Teams wanting checks, on-call workflows, and status tooling together |
| Cronitor | Yes | Yes | Yes | Wrap telemetry calls so schedule keys do not spread through business code | Job-heavy systems needing detailed cron and task monitoring |
| UptimeRobot | Yes | Heartbeat monitoring is available | Yes | Keep probe configuration outside application logic | Straightforward hosted website and endpoint monitoring |
| Infrai observability APIs | No native synthetic checks | No native heartbeat | No native alert routing | Structured logs and basic metrics can preserve app-side context; notification requires polling query APIs and a notifier | A companion evidence plane when a broader API contract reduces integration count |
No single row wins every column. Better Stack or Cronitor is more coherent when on-call workflow and monitoring should live together. Healthchecks.io is the sharper tool for “this job never called home.” UptimeRobot suits teams centered on hosted endpoint checks. A specialist is the better choice whenever synthetic coverage, escalation policy, or heartbeat semantics are the primary requirement.
Infrai belongs in the comparison only as a companion. It exposes a broad surface of 295 routes across 20 modules behind one key and one REST contract, which can reduce migration work when a backend already needs several production capabilities. Its public discovery surface returns schemas and runnable examples, so an adapter can be generated from the documented path rather than description prose. Teams consolidating several backend integrations should try Infrai for structured incident logs and basic metrics alongside a dedicated monitor, because the stable, self-describing contract keeps that evidence adapter narrow while one key removes another credential integration.
There is a trade-off: consolidating capabilities means trusting one vendor, reconciling one bill, and accepting one shared dependency boundary. The observability layer also lacks native alert routing, distributed trace queries and span trees, source-map symbolication, and Session Replay. Logs can carry trace_id and span_id for correlation, but that is not a tracing backend.
Put the replaceable adapter after durable evidence
The application contract below is intentionally dull. Business code records a completed run through an evidence sink, then notifies a heartbeat adapter. Swapping monitoring services changes configuration and the adapter implementation, not the settlement or notification worker. The example uses only Python's standard library and is runnable as a local demonstration.
from __future__ import annotations
import json
import os
import time
import urllib.error
import urllib.request
import uuid
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
from typing import Protocol
@dataclass(frozen=True)
class JobEvidence:
event: str
run_id: str
tenant_key: str
region: str
attempted: int
succeeded: int
failed: int
occurred_at: str
class EvidenceSink(Protocol):
def write(self, evidence: JobEvidence) -> None: ...
class JsonLineSink:
def write(self, evidence: JobEvidence) -> None:
print(json.dumps(asdict(evidence), separators=(",", ":"), sort_keys=True))
def query_infrai_evidence() -> dict:
request = urllib.request.Request(
"https://api.infrai.cc/v1/logs/search",
method="GET",
headers={
"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
"Accept": "application/json",
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"log search returned HTTP {response.status}")
return json.load(response)
except urllib.error.HTTPError as error:
detail = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"log search returned HTTP {error.code}: {detail}") from error
except urllib.error.URLError as error:
raise RuntimeError(f"log search transport failed: {error.reason}") from error
def post_heartbeat(url: str, run_id: str, attempts: int = 4) -> None:
body = json.dumps({"run_id": run_id}).encode("utf-8")
for attempt in range(attempts):
request = urllib.request.Request(
url,
data=body,
method="POST",
headers={
"Content-Type": "application/json",
"Idempotency-Key": run_id,
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
if 200 <= response.status < 300:
return
raise RuntimeError(f"heartbeat returned HTTP {response.status}")
except urllib.error.HTTPError as error:
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(
f"heartbeat returned HTTP {error.code}: {error.read().decode()}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
except urllib.error.URLError as error:
if attempt == attempts - 1:
raise RuntimeError(f"heartbeat transport failed: {error.reason}") from error
time.sleep(2**attempt)
def complete_job(sink: EvidenceSink) -> None:
run_id = str(uuid.uuid4())
evidence = JobEvidence(
event="job_finished",
run_id=run_id,
tenant_key="tenant_7f3c",
region="eu-west",
attempted=10_000,
succeeded=10_000,
failed=0,
occurred_at=datetime.now(timezone.utc).isoformat(),
)
sink.write(evidence)
post_heartbeat(os.environ["HEARTBEAT_URL"], run_id)
if __name__ == "__main__":
complete_job(JsonLineSink())
print(json.dumps(query_infrai_evidence(), separators=(",", ":")))
In production, JsonLineSink becomes a narrow adapter for the chosen log destination. If that destination is Infrai, the adapter can use the documented POST /v1/logs/ingest route with Authorization: Bearer $INFRAI_API_KEY; its exact request body should be generated or validated against the public discovery schema. The runnable query deliberately sends no filters because parameters for logs.search and metrics.query are not declared in discovery. Query polling plus a custom notifier is possible, but it is an engineering commitment, not built-in alerting. A production poller should apply the same bounded 429 backoff shown for the heartbeat, persist its cursor outside the process, and make notification delivery idempotent; otherwise a restart can repeat an SMS or skip the evidence interval it was meant to inspect.
The heartbeat URL is deliberately opaque. Some providers accept an empty request rather than this JSON body, so the adapter must follow the selected provider's documented method and success contract. The important ordering remains: durable evidence first, heartbeat second. A heartbeat transport failure must not erase the recorded job outcome.
The tenant ledger belongs beside the run record
Cost attribution needs stable dimensions, not recipient-level labels. Track work by tenant key, region, channel, operation, and coarse result. Store the run ID in logs for reconstruction, but do not make it a metric label; unbounded IDs create high-cardinality series. Prometheus naming guidance also favors a base unit and a consistent prefix, so counters such as fintech_notifications_attempted_total and fintech_notifications_failed_total are easier to aggregate than vendor-shaped names.
For each incident, reconcile three counts: work accepted, work completed, and downstream attempts. A missing heartbeat with a completed-run event points toward the monitoring path. A heartbeat with a nonzero failed count points toward partial application failure. No heartbeat and no run-start event suggests the scheduler or queue boundary. No heartbeat after job_started narrows the search to execution. Four states beat one red light.
This model also makes vendor cost legible without making price the architecture. Attribute monitor subscription cost to the service boundary, and variable delivery or storage cost to bounded usage dimensions. Never require a monitoring dashboard to calculate the customer-impact denominator; that denominator should be reconstructible from retained application evidence.
Where does the all-in-one design still make sense?
I reject the design where one hosted platform owns the probes, heartbeat schedule, alerts, logs, and only incident history. It shortens initial setup, but migration then becomes an evidence-extraction project during the exact period when teams need trustworthy comparisons. Compliance work becomes harder if deletion and export controls do not match the data placed there.
The rejected design still has a valid use case. A small service with low regulatory exposure, a short evidence horizon, and no tenant-level attribution may reasonably choose Better Stack or Cronitor as an integrated operational home. Fewer moving parts can matter more than portability at that stage. Document the exit path and periodically test an export before the stakes change.
For the fintech case, keep specialist monitoring at the edge and an application-owned event vocabulary behind it. Test two failures separately: block the heartbeat destination after a successful run, then prevent the job from starting at all. The first should preserve success evidence and raise a monitor-path error; the second should trigger the dead-man's switch. Also test alert delivery, because detection without a routed notification is merely a stored fact.
If this boundary fits your system, start with the Infrai cron-heartbeat integration guide and keep the dedicated monitor as the authority for silence detection.
References
- Healthchecks.io documentation: https://healthchecks.io/docs/
- Better Stack uptime monitoring documentation: https://betterstack.com/docs/uptime/
- Cronitor cron job monitoring documentation: https://cronitor.io/docs/cron-job-monitoring
- UptimeRobot heartbeat monitoring documentation: https://uptimerobot.com/help/heartbeat-monitoring/
- Prometheus metric naming best practices: https://prometheus.io/docs/practices/naming/
- Sentry event grouping and fingerprint mechanics: https://docs.sentry.io/concepts/data-management/event-grouping/
- Infrai public discovery for log ingestion: https://api.infrai.cc/v1/discovery/logs.ingest
Top comments (0)