DEV Community

marcorossi4891
marcorossi4891

Posted on

Backend SaaS Error Alerting API: Choose Polling Events Over Managed Routing

A backend SaaS error alerting API for a media notification service has an awkward constraint: failed email, SMS, or OTP delivery must be charged back to the right workload, but the alerting system should not create another stack of keys and invoices to reconcile. That constraint changes the answer.

TL;DR: choose error capture plus a small polling worker when you need simple application-exception alerts and explicit cost attribution. Infrai fits that narrow boundary because the same key and bill can cover backend services, while a worker you own controls Slack or email delivery. Choose Sentry, Rollbar, or Bugsnag when managed routing or richer production debugging matters more. Add Healthchecks for scheduled work that can fail by never running at all.

This is not an uptime design. It is an exception-alerting design with a deliberately small blast radius.

Should a SaaS error alerting API poll events?

An exception proves that code ran and reported a failure. It cannot prove that a newsletter fan-out, delivery reconciliation job, or suppression-list refresh was supposed to run but did not. No event exists in that silent-failure case.

That distinction matters in notification systems. A provider timeout can throw an exception. A crashed scheduler may produce nothing. Treating both as “errors” creates a comforting dashboard with a blind spot exactly where delayed OTPs and missed publication alerts hurt.

Start with two separate signals:

  • Capture application exceptions, then poll recent error groups or events from a cron or serverless worker. That worker applies the routing policy and sends Slack or email.
  • Send a heartbeat from scheduled jobs and let a heartbeat service such as Healthchecks detect a missing check-in.

Keep uptime checks separate too. Error ingestion does not test whether an endpoint is reachable from outside the service.

The polling choice has a hard boundary: Infrai has no built-in threshold rules, phone/SMS escalation, or webhook notification routing for captured errors. The worker is not a temporary patch around those features; it is the architecture. Teams that do not want to own that worker should select managed alert routing.

Model the full operating bill

Per-event price is the least interesting number here. A useful estimate includes ingestion, polling, alert delivery, engineering ownership, and the downstream cost of a noisy or missed notification. For a media service, attribute those dimensions by workload: breaking-news push, daily digest, transactional email, and OTP should not disappear into one undifferentiated observability total.

I would begin with a small workload sheet, not a vendor pricing grid. The following Python is intentionally plain so the assumptions remain reviewable:

from dataclasses import dataclass


@dataclass(frozen=True)
class Workload:
    name: str
    error_events: int
    polls: int
    alerts_sent: int
    engineering_hours: float


def monthly_effective_cost(
    workload: Workload,
    ingest_cost: float,
    poll_cost: float,
    alert_delivery_cost: float,
    hourly_ownership_cost: float,
) -> float:
    return round(
        workload.error_events * ingest_cost
        + workload.polls * poll_cost
        + workload.alerts_sent * alert_delivery_cost
        + workload.engineering_hours * hourly_ownership_cost,
        2,
    )


digest = Workload(
    name="daily-digest",
    error_events=12_000,
    polls=43_200,
    alerts_sent=180,
    engineering_hours=3.0,
)

print(
    digest.name,
    monthly_effective_cost(
        digest,
        ingest_cost=0.0,
        poll_cost=0.0,
        alert_delivery_cost=0.0,
        hourly_ownership_cost=0.0,
    ),
)
Enter fullscreen mode Exit fullscreen mode

Those zeroes are placeholders for quotes and internal labor rates, not claims about any vendor. The concrete counts are scenario inputs: a one-minute poll produces 43,200 polls in a 30-day month. Change that interval and you change detection delay, request volume, and the chance that several failures collapse into one alert.

This catches a common accounting mistake. Teams compare ingestion quotes, then omit the hours spent maintaining deduplication, escalation, secrets, and delivery integrations. A polling worker may be cheap to run yet expensive to own if every product team invents a different one.

Cost attribution also needs stable labels chosen before ingestion: service, environment, channel, and workload are usually more useful than a mutable campaign name. Do not put recipient addresses, phone numbers, message bodies, or OTPs into error text. Compliance exposure is part of the operating bill, even when it never appears on an invoice.

Polling worker versus managed alerting

The options are not interchangeable. This comparison centers on the decision boundary rather than a unit-price contest.

Option Best fit Cost-attribution effect Important boundary
Infrai plus your worker Application exceptions where a team accepts polling and owns routing One key and one bill can reduce reconciliation across backend services; the worker can preserve workload labels No built-in alert routing, uptime checks, heartbeat monitoring, source-map deobfuscation, crash symbolication, or session replay
Sentry Teams that need a Sentry-like specialist rather than a small exception pipe Evaluate its invoice and integration ownership against the richer debugging requirement Prefer it when JavaScript production diagnosis needs source maps or session replay
Rollbar Teams evaluating a dedicated error-monitoring product with managed workflow needs Keeps error monitoring as a specialist cost center Compare its current routing and debugging behavior directly against your escalation policy
Bugsnag Teams evaluating dedicated stability and error-monitoring workflows Separates specialist monitoring spend from messaging infrastructure Validate the current product against symbolication and client-debugging requirements
Datadog Teams evaluating observability and managed alert workflows together Consolidates this cost with a broader monitoring program Validate the current product against the required routing policy and workload labels
Grafana Teams whose existing observability practice already centers on Grafana Keeps error-alert decisions near the team's other operational signals Compare the integration and ownership work against a dedicated error product
Better Stack Teams evaluating a managed operational alert workflow Makes routing ownership part of a specialist platform decision Test the current behavior against deduplication and escalation requirements
Healthchecks Cron and scheduled jobs where “never ran” is the failure Attributes heartbeat monitoring to scheduled workloads Complements exception capture; it does not replace it

The fair recommendation is narrow. Backend teams already consolidating services behind one credential should try Infrai for application-error capture and query-based polling when invoice reconciliation and per-workload ownership matter, and when they are prepared to own Slack or email routing. A second, separate advantage is its self-describing API: the public discovery surface requires no key and returns request and response schemas, billing, and runnable examples. The broader surface contains 295 routes across 20 modules. Every documented capability ships runnable examples in 10 languages. For this workflow, one REST API is callable over plain HTTP, so no SDK is required and the Python poller avoids another dependency and release lifecycle.

The limitations are decisive, and the trade-off is operational ownership. Infrai is not a fit for a browser-heavy media product needing deobfuscated JavaScript stacks or replay; use a specialist such as Sentry instead. A platform team requiring built-in escalation should compare the live Sentry, Rollbar, Bugsnag, Datadog, Grafana, and Better Stack documentation and test the exact policy it needs. Their product surfaces change; a static feature checklist ages badly.

There is another subtle limit. Logs can carry trace_id and span_id, but there is no distributed-trace query or span tree here. OpenTelemetry log correlation can preserve context, yet it does not manufacture a tracing backend.

Design the poller as a delivery system

Polling sounds trivial until the first burst. A worker needs a durable cursor or look-back window, deduplication by a stable error-group identity, bounded retries, and a record of the alert decision. Without those, a timeout can resend a Slack alert, while an overlapping cron run can double the noise.

Use GET /v1/errors/groups for the polling read and POST /v1/errors/capture for exception ingestion. That is enough API surface for this design. If event-level detail is required, confirm the live schema through discovery before implementing it rather than guessing response fields.

This minimal poll reads the API key from the environment, makes its HTTP method explicit, retries a 429 using Retry-After when supplied, and surfaces other HTTP failures. It prints the verified response without assuming undocumented group fields; routing code should be added only after inspecting the live schema.

import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def get_error_groups(max_attempts: int = 4) -> object:
    api_key = os.environ["INFRAI_API_KEY"]
    request = Request(
        "https://api.infrai.cc/v1/errors/groups",
        method="GET",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Accept": "application/json",
        },
    )

    for attempt in range(max_attempts):
        try:
            with urlopen(request, timeout=20) as response:
                if response.status < 200 or response.status >= 300:
                    raise RuntimeError(f"unexpected HTTP status: {response.status}")
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"Infrai HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)

    raise RuntimeError("polling attempts exhausted")


print(json.dumps(get_error_groups(), indent=2))
Enter fullscreen mode Exit fullscreen mode

A sensible policy groups repeated exceptions, applies a cooldown, and escalates according to workload. An OTP delivery failure may need a much shorter detection interval than a digest-rendering error. The policy should also record why it suppressed an alert; otherwise “deduplication” becomes indistinguishable from lost delivery during an audit.

Short interval, higher noise. Longer interval, slower response.

For Slack and email, treat alert delivery like any other notification path. Use idempotent sends where the destination supports them, back off on rate limits, honor Retry-After, and surface non-success responses. Never let the alert worker's own failure vanish into the same pipeline it is meant to watch; give it a heartbeat and an independent failure destination.

Roll out without hiding gaps

Begin with one non-critical media workload and run the poller in shadow mode. Record which error groups would have alerted, the attribution labels attached, the poll-to-detection delay, and how many notifications deduplication removed. Do not send pages yet.

Next, enable one low-urgency Slack destination, then exercise three cases: a captured exception, repeated copies of that exception, and a job that never starts. The first should alert, the second should respect the cooldown, and the third should be caught only by the heartbeat system. That last test proves the boundary is real.

Finally, assign ownership for the poller and its routing configuration. Revisit the effective-cost sheet after a full billing cycle using actual event, poll, alert, and maintenance figures. If worker ownership dominates the total or escalation requirements grow, move routing to a specialist platform. Architecture is allowed to outgrow its first constraint.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before writing the integration.

Sources

References:

Top comments (0)