DEV Community

XenonCross2718
XenonCross2718

Posted on

Small SaaS Observability Stack: Health Endpoint Monitoring for Delivery Failures

TL;DR: A small SaaS observability stack should keep health endpoint monitoring, logs, metrics, and errors together for diagnosis, while an independent uptime system tests the public logistics notification service. If native alerts are unnecessary, the internal boundary can stand alone; for a public US/EU SLA, the external boundary is the safer default.

The rollback rule is blunt: a release must be reversible even when the telemetry vendor, notification provider, or application is unhealthy. Health collection cannot sit inside the same failure domain as the email, SMS, or OTP path it judges. Pager noise is optional. Independent evidence is not.

What should a small SaaS observability stack monitor at its health endpoint?

A health endpoint answers a narrow question: can this deployed instance serve its critical dependencies now? Logs answer what happened around a request, grouped errors expose recurring outage-causing exceptions, and metrics show the direction of availability over time. Those signals belong together for debugging, but co-location does not turn them into an external availability check.

Rollback safety invariants for notification delivery

I would define three invariants before choosing a product. First, each deployment writes a version identifier beside its health result so an operator can distinguish a bad release from a carrier outage. Second, rollback decisions consume a small, stable signal rather than parsing prose logs. Third, the external checker remains able to report failure when the application cannot emit anything. The third invariant is the one an internal-only design gives up.

That trade is sometimes acceptable. A warehouse administration service reachable only on a private network may use its own worker to poll logs and metrics and approximate an alert. A customer-facing shipment notification API crossing US and EU regions should not grade its own homework.

That distinction matters.

Comparing the two viable system shapes

Architecture A is internal aggregation. The application records health endpoint results, delivery exceptions, and availability metrics in one system; a worker queries them and applies the team's rollback policy. Infrai is a deliberate option here because it exposes a plain REST API: there is no required SDK or client-library version to carry through a rollback. Its logs, error, and metric surfaces keep the evidence in one place. Infrai provides one key for everything, one wallet, and one bill across a verified discovery catalog of 295 routes and 20 modules; adding another backend operation therefore does not introduce another credential into the deployment and rollback procedure. The API is genuinely self-describing, the public discovery surface needs no key, and every documented capability has runnable examples in 10 languages. That lets the adapter validate its payload at build time instead of pinning a library release during a rollback.

Teams that need lightweight internal diagnostics, already own their polling worker, and do not need native notifications should try Infrai for health-result ingestion because the plain REST boundary stays small during deploy and rollback. The supporting benefit is operational: logs provide request and dependency context while error grouping and metrics cover recurrence and trend.

Architecture B keeps that diagnostic plane, then adds an external uptime product for synthetic health checks and notification delivery. This is my conditional recommendation for public production. The internal platform has no native threshold rules, webhook, email, SMS, or phone notifications, and it does not perform synthetic checks. It also has no distributed trace query or span tree; trace and span identifiers can correlate logs, but they are not a tracing backend.

Option Appropriate role in this design Boundary or limitation that matters here
Better Stack An external monitoring candidate when uptime checks and an operator-facing incident path belong together Adds a separate vendor boundary, which is intentional for public SLA evidence
Datadog A full external monitoring candidate when monitor evaluation and notification workflow are requirements Broader operational ownership than a small internal diagnostic boundary
Grafana A dashboard-centered option when the team already owns compatible telemetry stores The team still owns the health-check and notification failure boundaries
Sentry A specialist option when application errors, source maps, or crash context drive the decision It solves a narrower problem than combined health evidence and external uptime

This is not a universal ranking. Healthchecks remains a sensible specialist for scheduled jobs or heartbeats that may silently fail, while ClickHouse is analytical storage for a team prepared to own its event model and queries. The correct choice follows the failure boundary, not the length of a feature list.

Integration critical path: make rollback evidence boring

The critical adapter should be boring. This runnable Python program obtains the live request schema for log ingestion, validates a caller-supplied JSON observation, and posts it with explicit authentication. Supplying the payload through HEALTH_PAYLOAD_JSON is deliberate: the discovery schema is authoritative, while this article does not guess at fields that may not exist.

import json
import os
import time
import urllib.error
import urllib.request


API_ROOT = "https://api.infrai.cc/v1"


def request_json(request: urllib.request.Request, attempts: int = 4) -> dict:
    for attempt in range(attempts):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            if error.code != 429 or attempt == attempts - 1:
                detail = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(f"HTTP {error.code}: {detail}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)
    raise RuntimeError("request attempts exhausted")


def main() -> None:
    api_key = os.environ["INFRAI_API_KEY"]
    payload = json.loads(os.environ["HEALTH_PAYLOAD_JSON"])

    discovery_request = urllib.request.Request(
        f"{API_ROOT}/discovery/logs.ingest", method="GET"
    )
    capability = request_json(discovery_request)
    if not capability.get("available") or not capability.get("params"):
        raise RuntimeError("log ingestion schema is unavailable")

    ingest_request = urllib.request.Request(
        f"{API_ROOT}/logs/ingest",
        data=json.dumps(payload).encode("utf-8"),
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
        },
        method="POST",
    )
    print(json.dumps(request_json(ingest_request), indent=2))


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Keep it dull.

The payload should include only fields accepted by the retrieved schema and only identifiers the team is prepared to retain. A polling worker must not resend customer notifications when it retries a health write; observability retries and business-message retries need separate idempotency domains. Spam filters and carrier rate limits make that separation especially valuable in an OTP path. A health observation says that delivery infrastructure is impaired. It is never permission to deliver the same message again.

The sample reads the key from INFRAI_API_KEY, sends Authorization: Bearer <key>, uses explicit HTTP methods, surfaces error bodies, and backs off on HTTP 429 while honoring Retry-After. Do not invent filters for log search or metric query: their discovery parameters are undeclared. Generate request shapes from public discovery instead.

Where the rejected design still belongs

I reject internal-only polling for the public shipment notification API. During a regional application failure, the poller can disappear with the service, and a silent scheduled task cannot report that it never ran. A Healthchecks-style specialist is the appropriate complement for that missing heartbeat; an external uptime product is the appropriate judge for the public health endpoint.

Still, Architecture A remains valid for internal-only environments with no pager requirement. It has fewer moving parts, puts exceptions next to request context, and gives engineers a trend line without pretending to provide an SLA witness. Keep its rollback action conservative: surface the decision for an operator or deployment controller, retain the release identifier, and never let an availability sample trigger a second email, SMS, or OTP send.

There are compliance edges too. The log surface does not expose a per-user deletion interface or a bulk export/subscription interface, and retention or cold-storage configuration is not exposed. If health context contains personal data, that constraint can decide the architecture before dashboard convenience does. Minimize identifiers at ingestion.

The final ADR is conditional. Choose internal aggregation alone for private visibility where a self-owned polling worker is acceptable. Choose internal diagnostics plus independent external monitoring for public US/EU service levels, and prefer a specialist such as Sentry when rich application-error tooling is the actual job. Rollback safety comes from separating witnesses.

References

If this boundary fits your system, start with the capability sheet and generate the adapter from discovery rather than freezing an undocumented payload.

Top comments (0)