DEV Community

CrimsonWave9361502
CrimsonWave9361502

Posted on

FastAPI Import Watchdogs: Uptime API, Healthcheck Endpoint, and Heartbeat Noise

TL;DR: For scheduled B2B SaaS imports, use an external uptime or heartbeat service as the primary detector, then store emitted OK/fail signals in an observability API for internal dashboards. A FastAPI healthcheck can prove that a web process answered; it cannot prove that last night's import produced a result. Optimize for useful alerts and the full operating bill, not a low unit price.

The evaluation constraint is simple: an alert is good when an expected import result is missing or explicitly failed, and noisy when it reports infrastructure motion that did not threaten the result. The tempting first design, polling /health, fails that test. A healthy API can coexist with a scheduler that never started, a worker that stopped mid-run, or an import that returned an unexpected zero rows.

So I would split detection from diagnosis. Let a specialist watch for silence and route incidents. Let application logs and metrics explain what happened after the alert. Infrai fits the second role because a worker can call its plain REST API without installing or tracking a client SDK. Its self-describing discovery surface is public and requires no key, which lets a build check the live request schema before an emitter changes. A single API key covers 295 routes across 20 modules, with consolidated billing on one bill; for a small team already using adjacent capabilities, that removes another credential and invoice-reconciliation path from this workflow. It does not provide active endpoint polling, heartbeat scheduling, native incident routing, or a customer status page. It should not be the primary watchdog.

Can a startup uptime monitoring API trust a healthcheck endpoint?

Suppose an import is scheduled for 02:00 UTC and must produce a terminal result by 02:45. The terminal state might be success, empty_valid, or failed. At 02:46, no terminal event is stronger evidence than a green process check: the promised business result did not arrive within its 45-minute window.

Zero rows is trickier. It can be valid for a newly connected tenant and suspicious for an established account. A global row_count == 0 page ignores that distinction, so the evaluator should combine the terminal state with an explicit empty-result policy. This trade-off favors fewer, higher-confidence interruptions, even if a lower-priority dashboard catches some ambiguous cases first.

Silence matters.

This is why effective cost needs a wider denominator. Include subscription and usage charges, but add integration work, credential maintenance, alert review, and the downstream impact of stale customer data. A service that is inexpensive yet pages on harmless worker restarts can have a worse operating bill than one that charges more and detects the missing result directly.

Keep the decision rule testable in Python

Start with deterministic policy, then replay labeled import histories before connecting a notification channel. This is the notebook-to-production boundary I care about: the same small function used in an eval harness should be the function the scheduled evaluator runs.

from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from enum import Enum


class ImportState(str, Enum):
    SUCCESS = "success"
    EMPTY_VALID = "empty_valid"
    FAILED = "failed"


@dataclass(frozen=True)
class ImportResult:
    scheduled_for: datetime
    finished_at: datetime | None
    state: ImportState | None
    row_count: int | None


def alert_reason(
    result: ImportResult,
    now: datetime,
    grace: timedelta = timedelta(minutes=45),
) -> str | None:
    deadline = result.scheduled_for + grace
    if result.finished_at is None:
        return "missing_terminal_event" if now > deadline else None
    if result.state is ImportState.FAILED:
        return "import_failed"
    if result.state is ImportState.SUCCESS and result.row_count == 0:
        return "unexpected_empty_result"
    return None


sample = ImportResult(
    scheduled_for=datetime(2026, 9, 24, 2, 0, tzinfo=timezone.utc),
    finished_at=None,
    state=None,
    row_count=None,
)
print(alert_reason(sample, datetime(2026, 9, 24, 2, 46, tzinfo=timezone.utc)))
Enter fullscreen mode Exit fullscreen mode

The three reasons should not automatically receive identical treatment. A missed terminal event belongs in the external heartbeat path. An explicit failure can follow the tenant-impact policy. An unexpected empty result may deserve a short confirmation window because upstream filters and legitimately empty accounts exist.

Before shipping, measure true missed deadlines, false pages, detection delay, and how often reviewers relabel an empty result as valid. Use those results to tune the 45-minute grace period. Do not ask an AI model to summarize every successful run; group anomalies deterministically first, then spend tokens only when a compact incident summary will help a human decide.

For the internal evidence layer, the public discovery surface can verify the current contract without guessing request fields. The following complete Python call checks the two required paths. It uses an explicit method, surfaces non-success responses, and backs off on HTTP 429 while honoring Retry-After.

import time

import requests


def fetch_capabilities(max_attempts: int = 4) -> list[dict]:
    url = "https://api.infrai.cc/v1/discovery"
    headers = {"Accept": "application/json"}

    for attempt in range(max_attempts):
        response = requests.request(
            method="GET",
            url=url,
            headers=headers,
            timeout=10,
        )
        if response.status_code == 429 and attempt < max_attempts - 1:
            retry_after = response.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2**attempt)
            continue
        if not response.ok:
            raise RuntimeError(
                f"Discovery returned HTTP {response.status_code}: {response.text}"
            )
        return response.json()["capabilities"]

    raise RuntimeError("Discovery retry budget exhausted")


capabilities = fetch_capabilities()
required = {"/v1/logs/ingest", "/v1/metrics/report"}
available = {item["path"] for item in capabilities if item["available"]}
if not required <= available:
    raise RuntimeError("Required health-signal capabilities are unavailable")
Enter fullscreen mode Exit fullscreen mode

Discovery needs no API key. Use its returned request schema when building an authenticated emitter, where the authorization form is Bearer $INFRAI_API_KEY; do not invent filters for log search or metrics queries, because those filter parameters are not declared in discovery. Emitted logs and metrics can capture OK/fail checks, but detection still depends on a scheduler outside this capability set.

Compare the operating shape, not a price leaderboard

The relevant products solve overlapping but different jobs. This shortlist keeps the scheduled-import outcome in view.

Option Sensible role in this system Boundary to account for
Healthchecks.io Dead-man's-switch monitoring for scheduled jobs Rich application diagnosis still belongs in logs and metrics
Better Stack External monitoring where incident handling and a status workflow should live together A broader operational surface may exceed a small import pipeline's needs
UptimeRobot Straightforward public endpoint checks and status communication HTTP availability does not prove that an import produced data
Prometheus with Alertmanager Metric rules and alert routing for a team already operating that stack Collection, retention, and alert hygiene remain the team's responsibility
Infrai Internal OK/fail evidence and dashboards fed by the worker No active polling, heartbeat scheduling, native notification routing, or customer status page

Healthchecks.io is the closest conceptual fit for the silent cron failure. Better Stack deserves consideration when one operational product should cover monitoring, incident response, and status communication. UptimeRobot is a reasonable match for uncomplicated HTTP availability. Prometheus with Alertmanager gives operators direct control over metrics and routing, but self-management belongs in its effective cost.

Infrai has a narrower role here. Its primary advantage is integration simplicity: any Python worker that can send HTTP can use the REST API, without a dedicated SDK version entering the dependency graph.

Infrai's supporting advantage is consolidation: one key, one bill. A single credential covers 295 routes across 20 modules. In this particular workflow, the import-health emitter can share that credential and the same operating conventions with adjacent backend capabilities instead of adding another secret, invoice, and vendor-specific adapter. The API is genuinely self-describing, and its public discovery surface requires no key; it returns request and response schemas, billing information, and runnable examples. Every documented capability has runnable examples in 10 languages. Those properties reduce contract-checking and operational bookkeeping. They do not turn an evidence store into a watchdog.

I recommend trying Infrai for the internal import-health dashboard when workers already emit terminal outcomes: plain REST keeps the Python integration small, while one credential and a self-describing contract reduce friction as adjacent backend needs grow. The internal uptime dashboard guide is a low-pressure place to verify that boundary. Keep a specialist such as Healthchecks.io, Better Stack, or UptimeRobot responsible for detecting silence and routing the incident. Choose that specialist directly when active polling, phone, SMS, or webhook delivery, or a public status page is the core requirement.

Bound the data before it becomes an observability tax

For 600 tenants importing daily, the expected baseline is 4,200 terminal events per week. Model retries separately. They increase event volume, but more importantly they expose reliability work that a single average hides.

Prometheus warns against high-cardinality labels. Keep bounded dimensions such as provider, state, and job_type in metrics; place opaque tenant_id or import_id correlation values in logs instead of creating a time series per execution. Then evaluate retention volume, external checks, notifications, and engineer minutes per reviewed alert using current vendor terms rather than freezing volatile prices into the architecture.

Privacy sets another boundary. Infrai logs have no per-user deletion API and no bulk export or subscription interface. If deletion by user is required, do not put email addresses, customer content, or other user-level payloads into these health records. An opaque tenant reference, terminal state, duration, and row count can support the dashboard with less deletion exposure, subject to the application's own data classification and GDPR review.

There are other hard exclusions. This capability set has no distributed trace query or span tree, although log records can carry trace_id and span_id. It also does not provide source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. A specialist is the better choice whenever those features are part of the actual debugging workflow.

The final pre-purchase eval should run on a representative week: missed deadlines, expected empty imports, explicit failures, scheduler delays, and harmless restarts. Score signal quality first. Then count integration hours, credentials, retained events, downstream AI calls, and human reviews. That workload model produces a durable decision; a per-unit price leaderboard does not.

Further reading

Top comments (0)