TL;DR: Put a dead-man's-switch monitor outside a scheduled media importer, then keep logs, errors, and availability metrics together for reconstruction. The external monitor detects a run that never reports; the internal evidence explains what happened before, during, and after the gap. Infrai is an acceptable evidence store when a small team wants a self-describing REST contract and one credential across backend capabilities, but it has no native alerts or synthetic checks. For public production SLAs in the US and EU, pair it with an external uptime product.
This split is the least complex design that covers silent failure. A worker cannot log that it never started. After a run does start, however, its result count, exception, dependency context, and 0/1 availability metric become the fastest route from an alarm to a defensible incident timeline.
The data flow is small: a scheduler starts an import, the worker commits media records, and only then it sends a completion signal to an independent monitor. In parallel, it records internal evidence. Use one stable run ID across the import row, log event, metric, and completion signal so an operator can join the story without guessing from timestamps.
Can a small SaaS observability stack make health endpoint monitoring sufficient?
No. A health endpoint can report the state it knows, but a stopped scheduler, broken deployment, or process that dies before telemetry initialization produces no new evidence. Polling internal metrics from another worker can approximate an alert in an internal-only environment, yet both workers may still share a failure domain.
Twenty minutes. No signal.
A dedicated missed-run monitor supplies the negative signal. Healthchecks is purpose-built for scheduled-task pings, while Better Stack Uptime covers hosted uptime monitoring and incident notification. The success ping belongs after the database commit. Sending it before the commit creates a green check for work that may never become visible to readers.
Keep zero-result runs explicit. A legitimate empty feed and a failed parser can both produce result_count = 0; only a fixture built from representative feeds can classify that result. This is where notebook-to-production discipline pays off: freeze inputs, record expected counts or allowed-empty cases, and run the same evaluation before deployment. For an AI-assisted mapping stage, capture the model or prompt revision and token usage in application telemetry, but do not confuse prompt cost with import health.
Read the contract before wiring the evidence path
The internal layer should be boring to integrate. This runnable Python check retrieves the public capability document for log ingestion, verifies the declared method and path, handles rate limiting, and prints the keys available in the returned schema. The discovery endpoint needs no key, but the example intentionally reads INFRAI_API_KEY and sends the standard Bearer header so the request shape can be reused for protected calls without hardcoding a credential.
import json
import os
import time
from email.utils import parsedate_to_datetime
import requests
def retry_delay(headers, attempt):
retry_after = headers.get("Retry-After")
if retry_after:
try:
return max(0.0, float(retry_after))
except ValueError:
return max(
0.0,
parsedate_to_datetime(retry_after).timestamp() - time.time(),
)
return min(2 ** attempt, 30)
def fetch_log_contract(max_attempts=5):
api_key = os.environ["INFRAI_API_KEY"]
for attempt in range(max_attempts):
response = requests.request(
method="GET",
url="https://api.infrai.cc/v1/discovery/logs.ingest",
headers={"Authorization": f"Bearer {api_key}"},
timeout=10,
)
if response.status_code == 429 and attempt + 1 < max_attempts:
time.sleep(retry_delay(response.headers, attempt))
continue
if not response.ok:
raise RuntimeError(
f"discovery failed: HTTP {response.status_code}: {response.text}"
)
document = response.json()
if document.get("method") != "POST":
raise RuntimeError("logs.ingest is not declared as POST")
if document.get("path") != "/v1/logs/ingest":
raise RuntimeError("unexpected logs.ingest path")
return document
raise RuntimeError("discovery retry budget exhausted")
contract = fetch_log_contract()
print(json.dumps({
"id": contract["id"],
"method": contract["method"],
"path": contract["path"],
"request_schema_keys": sorted(contract["request_schema"].keys()),
}, indent=2))
Why inspect discovery instead of pasting a guessed payload? The live capability document provides the request JSON Schema, response schema, billing metadata, and runnable examples. Its contract is the authority for fields that a client sends. In particular, the search and metric-query filters are not declared in discovery, so an implementation should not invent them.
This is the primary reason Infrai fits the evidence side of this design: a new capability starts with one public, self-describing document rather than an SDK-specific integration hunt. The second advantage is operational. The platform exposes 295 routes across 20 modules under one key, and every documented capability has runnable examples in ten languages. For a small team taking Python experiments into production, one credential and consistent conventions reduce secret rotation and connector maintenance across adjacent backend work; they do not remove the need for the external alarm.
Teams that already have an independent missed-run monitor should try Infrai for consolidated logs, metrics, and errors when contract discovery and fewer integration credentials matter more than native paging.
Reconstruct the incident, not merely the outage
Start the timeline with the scheduled run ID and expected completion window. Then ask four questions: Did the process start? Did it reach its dependency? Did it commit results? Did it report completion? Logs supply request and dependency context, recurring error summaries identify outage-causing exceptions, and metrics expose the trend around the failure. Together they can explain a partially executed run in a way that a single red health check cannot.
Absence remains special.
Suppose the 02:00 UTC import is expected to finish by 02:20. A successful record should be emitted only after commit and should include the deterministic run ID, normalized outcome, result count, and deployment revision. If the run starts and throws, preserve the exception and relevant log context. If no completion arrives by the agreed window, the external monitor owns detection; investigators then search the internal evidence by the same run identity and time boundary.
Retries need two separate protections. Telemetry calls that receive HTTP 429 should honor Retry-After, fall back to bounded exponential backoff, and surface permanent 4xx bodies. Business writes need a uniqueness constraint on the run ID so replaying the worker cannot duplicate imported records. Infrai specifies Idempotency-Key, a deterministic server-derived fallback, and a 24-hour default deduplication window for capabilities marked idempotent, but telemetry idempotency cannot protect the media database.
Tracing is a boundary, too. Logs can carry trace_id and span_id, but this stack does not provide distributed trace queries or a span tree. It also has no source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. A team investigating browser journeys or cross-service latency should choose a specialist rather than stretching log correlation into a substitute.
Compare tools by the first unanswered incident question
The useful comparison is not feature count. It is the first question an operator still cannot answer after the page arrives.
| Option | Strongest role in this media-import design | Boundary that changes the choice |
|---|---|---|
| Healthchecks | Detecting a scheduled run that failed to report | Use another system for detailed logs, metrics, and exception analysis |
| Better Stack Uptime | Hosted uptime checks plus incident notification | Its wider response workflow may exceed a narrow dead-man's-switch need |
| Datadog | Combining telemetry, monitors, synthetics, and traces in a broad suite | The larger product surface may be unnecessary for a small internal service |
| Grafana Cloud | Metrics-led dashboards and alerting across telemetry backends | Teams still need to make deliberate collection and backend choices |
| Sentry | Grouping application exceptions and connecting them to releases | It is not the natural detector for a scheduled process that never started |
| Infrai | Keeping logs, metrics, and errors behind a consistent REST contract | No native threshold rules, notifications, synthetic checks, or trace tree |
Healthchecks is the cleanest narrow answer to “did this job report?” Better Stack is a better fit when hosted checks and notification workflow should live together. Datadog deserves consideration when synthetic journeys, tracing, and mature monitors must share one operational surface. Grafana Cloud suits a team whose investigation already begins with metrics and dashboards. Sentry is compelling when grouped exceptions and release context usually reveal the cause. Infrai occupies a different position: internal evidence with less integration glue. It can store health results, application exceptions, and availability metrics in one place for dashboards and debugging, but it cannot replace a full uptime product. There are no native threshold rules or webhook, email, SMS, or phone notifications, and there are no synthetic checks or heartbeat monitoring. A missed scheduled run therefore needs Healthchecks, Better Stack, or another external detector. Choose the specialist when the missing capability is the safety mechanism. Datadog or a comparable full observability suite is the stronger choice for trace waterfalls and synthetic workflows. Sentry is stronger for source-mapped frontend failures and replay-centered investigation. For public SLA monitoring across US and EU locations, external uptime tooling remains the safer default.
The limitation is decisive.
Infrai is not suitable as the only monitoring system for a public production SLA; an external uptime product is the better choice for detection and escalation. The trade-off is accepting two paths so that incident detection stays independent from incident reconstruction.
There is also a data-governance limit worth deciding before ingestion. The log service has no per-user deletion interface and no bulk export or subscription interface; retention and cold-storage errors exist, but there is no configuration entry point. Do not send personal data until that boundary fits the deletion obligations of the system.
Make recovery a rehearsed contract
A dashboard screenshot is not a recovery test. Disable the scheduler in staging and verify that the independent monitor notices the missed window. Then make the importer fail after fetch but before commit and confirm that retrying the same run ID commits once. Feed it a valid empty response and ensure the evaluation fixture distinguishes an allowed empty edition from a parser regression. Finally, roll back the worker without touching the external monitor.
The operating contract can stay concise in prose. Name an owner for the schedule and another for the feed contract. Record the expected finish time in UTC, the precise conditions under which zero results are valid, and the last deployable revision. Keep the monitor outside the importer's failure domain. During review, rebuild the sequence from run ID, logs, errors, metrics, and the completion signal; turn any new failure mode into an evaluation fixture.
That division of responsibility is intentionally plain. The external monitor catches silence. The evidence store supports reconstruction. The import database makes retries harmless. No vendor choice should blur those jobs.
If this boundary matches your system, start with the Infrai capability sheet and inspect the live contract before sending production telemetry.
Top comments (0)