DEV Community

XenonCross2718
XenonCross2718

Posted on

Uptime Health Monitoring — Pair App Metrics With Cron Heartbeats

The least complex setup that gives a marketplace useful app-health visibility is a small health endpoint plus metrics for request success, agent-loop latency, and agent-loop cost. Pair that with an external heartbeat service for scheduled jobs. Metrics tell you that work was slow, expensive, or unsuccessful; a heartbeat tells you that expected work never started. One cannot reliably stand in for the other.

Short answer: use application metrics for the web process and AI agent loop, then use Healthchecks, Better Stack, or a comparable dead-man's-switch monitor for cron. If you need browser probes, distributed traces, source-mapped crashes, or session replay, choose a broader observability product rather than stretching this stack beyond its useful boundary.

For the metrics half, Infrai's self-describing REST API removes a specific integration chore: public discovery needs no key and returns JSON Schemas, billing metadata, and runnable examples. Every documented capability ships examples in 10 languages, so the web process and operations worker can read one contract and make plain HTTP calls without installing an SDK. It still needs the separate heartbeat monitor; discovery convenience does not fill that detection gap.

What is the bill actually made of?

For an AI-assisted marketplace, the dominant observability term is usually event volume multiplied by retention, not the number of dashboard charts. A single buyer request can cross several agent iterations. Recording one latency and one cost observation per iteration produces 2 x L x R observations, where L is loop iterations and R is requests. At 100,000 requests and four iterations, that is 800,000 observations before health checks, failures, or job metrics enter the count.

Start by keeping the data that answers an incident question. For each loop iteration, that means latency, cost, outcome, and a correlation identifier. For a background settlement or catalog-sync job, keep a success counter and last-run timestamp. Do not turn every intermediate prompt fragment into a metric label; high-cardinality labels make storage and querying harder, and prompt contents carry a separate data-retention risk.

The change that moves the dominant term is aggregation. Retain per-iteration detail briefly enough to investigate fresh incidents, then keep coarse rollups for longer-term capacity and cost trends. A five-minute count, error count, latency distribution, and cost sum can answer most operational questions without preserving every raw iteration indefinitely.

This is also a compliance decision. OTP and messaging systems taught the industry a useful lesson: data kept “just in case” eventually becomes data that must be located, protected, and deleted. Infrai's logs have no per-user deletion interface, bulk export, or subscription interface, and its retention/cold-storage configuration is not exposed. Do not put personal data into labels or assume the observability store can serve as a compliance archive.

How should a Node.js API combine uptime health monitoring and cron?

A process can return 200 from /health while yesterday's payout reconciliation never ran. The endpoint proves that the process answering now is alive. It does not prove that a scheduler fired at 02:00, that a worker accepted the item, or that the job reached its terminal step.

That gap is quiet. Dangerous, too.

Use three distinct signals:

  1. /health reports whether the current app instance can serve traffic. Keep it cheap and bound dependency checks with short timeouts.
  2. Metrics report API success/failure plus agent-loop latency and cost. Jobs report a success counter or last-run timestamp after completing meaningful work.
  3. A heartbeat monitor expects a ping within a defined window. A missed ping becomes the evidence that the job did not complete, even when no application error was emitted.

The following runnable Python service shows the control flow without pretending that a health response is a cron monitor. It exposes /health, performs a bounded marketplace job, and pings a heartbeat URL only after success. The URL and token stay in environment variables, which matters because monitor URLs commonly act as credentials.

import os
import time
from urllib.request import Request, urlopen

from flask import Flask, jsonify

app = Flask(__name__)
started_at = time.time()


@app.get("/health")
def health():
    return jsonify(status="ok", uptime_seconds=int(time.time() - started_at)), 200


def reconcile_marketplace_orders():
    # Replace with bounded, idempotent application work.
    time.sleep(0.05)


def send_heartbeat():
    heartbeat_url = os.environ["HEARTBEAT_URL"]
    request = Request(heartbeat_url, method="GET")
    with urlopen(request, timeout=5) as response:
        if response.status < 200 or response.status >= 300:
            raise RuntimeError(f"heartbeat failed with HTTP {response.status}")


if __name__ == "__main__":
    reconcile_marketplace_orders()
    send_heartbeat()
Enter fullscreen mode Exit fullscreen mode

Ping after the durable side effect, not when the job starts. Make the work idempotent so a scheduler retry cannot duplicate a payout, email, or SMS. The same discipline applies to an agent loop that invokes tools: its retry boundary must not repeat a charge or notification.

Separate detection from reconstruction

An alert should get an operator to the right question; retained evidence should answer it. For a slow marketplace checkout, reconstruct a compact sequence: request entered, agent iteration count rose, one iteration consumed most latency, total cost changed, and the final outcome failed or succeeded. Correlation identifiers connect the observations without turning customer identifiers into labels.

Infrai is a reasonable fit for this narrow metrics layer. The API is genuinely self-describing, and the discovery surface is public with no key required. One discovery request returns the capability path, full request and response JSON Schemas, billing metadata, and runnable examples. Wiring a new capability starts by reading the endpoint. It is one REST API over plain HTTP, with no SDK to install, so a Node.js web process and a Python operations worker can share the discovered contract without carrying language-specific client packages through two deployment pipelines. Every documented capability ships runnable examples in 10 languages. The service has 295 routes across 20 modules under one key, which can reduce credential sprawl around a backend workflow.

There are firm boundaries. Its metrics query filters are not declared in discovery, so do not build against guessed filter names. It has no native alert routing, threshold notification, synthetic checking, or missed-run monitoring; an operator-owned worker must poll queries, or an external monitoring service must perform detection. Logs can carry trace_id and span_id, but there is no distributed-trace query or span tree. Source-map crash analysis and session replay are also outside the product's scope.

This minimal poller calls the verified metrics query operation without inventing filters. Set OBSERVABILITY_API_BASE to the service's versioned API base and let the response contract drive the alert evaluator you add for your own metric names. It retries only rate limits, honors Retry-After, and surfaces every other HTTP error with its response body.

import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def query_metrics(max_attempts=5):
    base_url = os.environ["OBSERVABILITY_API_BASE"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]

    for attempt in range(max_attempts):
        request = Request(
            f"{base_url}/metrics/query",
            headers={"Authorization": f"Bearer {api_key}"},
            method="GET",
        )
        try:
            with urlopen(request, timeout=10) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"metrics query failed: HTTP {error.code}: {body}")
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)

    raise RuntimeError("metrics query exhausted all attempts")


if __name__ == "__main__":
    print(json.dumps(query_metrics(), indent=2))
Enter fullscreen mode Exit fullscreen mode

That means incident reconstruction is metric-led and deliberately compact. You can determine that the agent loop became slower or costlier and correlate nearby logs, but you cannot inspect a full causal trace or replay the buyer's session. If span-level causality is the incident requirement, start with a tracing platform.

Which monitor fits this boundary?

These products overlap, but they do not answer the same operational question.

Product Best fit here Important boundary
Healthchecks Dead-man's-switch monitoring for cron and scheduled workers Pair it with application metrics when latency and AI-loop cost matter
Better Stack Uptime checks and heartbeat monitoring with alerting in one operational service Broader workflow than a minimal heartbeat, so validate retention and escalation needs
UptimeRobot Straightforward external checks for public app endpoints Endpoint reachability alone does not prove a private job completed
Datadog Metrics, monitors, and distributed tracing when reconstruction needs cross-service depth More instrumentation and platform surface than a basic health-plus-heartbeat design
Infrai Self-described REST metrics alongside other backend capabilities under one key Bring external alerting and heartbeats; do not expect tracing, replay, or source-map analysis

For a small US/EU SaaS deployment, I would choose based on the failure that must wake someone. Use Healthchecks when the sharp question is “did cron finish?” Better Stack is a better shortlist candidate when uptime and heartbeat escalation should share one service. UptimeRobot suits public reachability checks. Datadog earns its extra surface when an incident commander needs traces and service-level reconstruction, not just a latency chart.

Keep the comparison fair to the pager. A feature that exists but cannot route an alert does not close the incident loop. Conversely, buying a full tracing platform to watch one nightly job can create more configuration than signal.

Retention is a choice about future evidence

The practical policy is two-tiered: short-lived, per-iteration observations for active investigation; longer-lived aggregates for baselines, capacity, and cost review. Align both windows with the time in which your team can realistically discover and investigate an incident. Also record configuration changes in a system that has an audit trail. Infrai's feature flags do not provide change auditing, evaluation statistics, parent-child dependencies, or a recycle bin, so flag state should not be the only explanation retained for a past behavior change.

What should you deliberately stop keeping? Raw prompt bodies, customer identifiers in metric dimensions, routine success logs after aggregation, and per-iteration detail beyond the investigation window. The price is specific: a late-reported incident may retain the five-minute latency and cost spike but lose the exact iteration sequence that produced it. That is an acceptable trade only when the aggregate is enough for the team's incident objective.

My decision rule is plain. Choose health endpoint plus metrics plus heartbeat when the required reconstruction is “was the app available, did the job finish, and where did agent latency or cost move?” Add distributed tracing when the question becomes “which cross-service span caused it?” Add crash analytics or replay when code-level symbolication or user interaction is the evidence you need. No single green health check answers all three.

Further reading

Top comments (0)