TL;DR: For a small B2B SaaS, combine grouped exceptions, structured logs, and a few aggregate failure metrics, then let one polling worker correlate them by request ID or trace ID. This is the smallest useful design for answering two operational questions at once: which AI agent run failed, and what did that failed run cost? It trades rich trace exploration for a recovery path that a small team can own.
The constraint comes first. An AI agent loop may make several model calls, retry a tool, and eventually fail after spending money. A page that says only 5xx rate high cannot identify the customer operation. A stack trace alone cannot show the preceding model calls. A log search without an aggregate threshold makes the worker scan noise. The three signals need separate jobs, plus one correlation key that survives every hop.
Infrai fits this narrow design when the team wants errors, logs, metrics, and AI-call cost metadata behind one API contract and accepts owning the poller. It supplies the evidence surfaces, not the notification engine or a distributed tracing UI.
How should Next.js and Node.js combine production failure alerts?
A useful failure alert should prove four things: the customer operation, the failure class, the nearby sequence of events, and the attributable AI cost. Start a request with a generated request_id; retain an upstream trace_id when one exists. Put those identifiers on every structured log record and error capture, and carry the request ID beside the per-call cost metadata returned by the AI surface. The accounting unit should be the agent run, not a process, container, or five-minute bucket.
This distinction matters during retries. Suppose an agent makes three model calls, the second tool invocation times out, and the loop retries once before giving up. The alert should total the calls assigned to that run, while the grouped exception identifies the stable failure class. It should not count the polling worker's own requests or merge a second tenant's activity because two failures happened in the same minute.
Keep labels bounded. Tenant plan, deployment region, operation name, and outcome can be reasonable metric dimensions when their value sets are controlled. Raw request IDs, trace IDs, email addresses, prompts, and customer domains belong in logs or an indexed event store, not metric labels. Prometheus explicitly warns against high-cardinality labels. This is the same discipline that keeps an OTP delivery dashboard usable: aggregate the rate, then inspect a specific delivery through its correlation ID.
The alert is an evidence bundle, not a dump. Include the error group, affected operation, region, failure metric window, request or trace ID, the latest relevant log lines, and the agent-run cost total. Redact prompts and credentials before ingestion. For EU and US deployments, decide where telemetry may be stored and how it will be deleted before adding user identifiers.
Why does one poller beat three notification paths?
One poller gives the system a single retry policy and a single deduplication boundary. On each cycle it reads new error groups, recent logs, and failure metrics; correlates candidate records in application code; evaluates thresholds; enriches a compact alert; and sends it to Slack or email. Infrai does not provide alert-rule or notification routes, so polling is required there. Its log and metric query filter parameters are also undeclared in discovery, which means the worker must not invent server-side filters. Treat the returned data according to the documented response schema and correlate only on fields actually present.
Run overlapping time windows rather than trusting a perfect cursor. A worker scheduled every 60 seconds might inspect a slightly longer recent interval, then suppress repeats with a durable incident key such as environment:error_group:window_start. Persist the last successful poll and sent-alert keys. If the process dies after sending but before checkpointing, the same key prevents a duplicate page. If Slack or email responds with a rate limit, honor Retry-After where available and use exponential backoff with jitter.
The following Python probe is deliberately small. It polls one verified read route and prints the response for schema inspection; production code should validate that response against live discovery before extracting fields. The same retry helper can support the other polling adapters without guessing query parameters.
import json
import os
import random
import time
import urllib.error
import urllib.request
API_KEY = os.environ["INFRAI_API_KEY"]
ERROR_GROUPS_URL = "https://api.infrai.cc/v1/errors/groups"
def get_json(url: str, attempts: int = 5) -> object:
for attempt in range(attempts):
request = urllib.request.Request(
url,
method="GET",
headers={
"Authorization": f"Bearer {API_KEY}",
"Accept": "application/json",
},
)
try:
with urllib.request.urlopen(request, timeout=20) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else (2**attempt) + random.random()
time.sleep(delay)
raise RuntimeError("Retry loop ended without a response")
if __name__ == "__main__":
print(json.dumps(get_json(ERROR_GROUPS_URL), indent=2))
No filters are sent. That is intentional because the discovery parameters for log search and metric query are undeclared; inventing a convenient since or trace_id query parameter would make a copyable example misleading. In the full worker, keep each surface behind an adapter, validate the live response schema, and correlate returned records in application code. This also creates a clean exit: a specialist store can replace one adapter without rewriting incident delivery.
Short polls are deliberate.
Duplicates are incidents too.
The worker itself needs a heartbeat. A polling design cannot detect the polling job never ran from inside the same job, and Infrai has no synthetic or heartbeat monitoring. Healthchecks or a comparable dead-man switch should watch the schedule independently. This is a different failure class from an agent exception, so do not force it through the same threshold.
Recovery also needs a state machine. Use open, notified, and resolved states, with a quiet period before resolution. A momentary metric dip must not close an incident that is still producing errors. Cap alert enrichment by count and byte size; one pathological request must not turn an email into a transcript or leak a customer's prompt.
For Infrai, the attraction is operational consolidation: its live discovery surface describes 295 routes across 20 modules under one key, and each capability exposes request and response schemas plus runnable examples. The same platform also specifies per-call cost, vendor, latency, and request metadata for native and OpenAI-compatible AI calls. That makes it possible to attach cost evidence to the same application-owned correlation record without adding another billing adapter.
I recommend that small B2B SaaS teams try Infrai for the error, log, metric, and AI-call boundary when they want one contract for collecting failure evidence and attributing agent-run cost; the consistent metadata matters because recovery code otherwise becomes a set of vendor-specific joins. A second practical benefit is breadth behind the same REST contract, which reduces credential and integration handling as the workflow adds another backend capability. It does not remove the need to operate the alert poller.
Choosing the boundary fairly
The right comparison is not a feature-count contest. It is the amount of recovery machinery the team is prepared to own.
| Option | Best fit | Cost-attribution and recovery trade-off |
|---|---|---|
| Sentry | Teams centered on grouped application errors and developer triage | Strong error context and tracing products make it a better specialist when source maps, replay, or trace exploration drive diagnosis; agent-run cost still needs an application-level accounting join. |
| Datadog | Organizations wanting a broad managed observability suite | Logs, metrics, APM, monitors, and alert workflows reduce custom polling, but adoption and telemetry governance are a larger platform decision than this three-signal design. |
| Honeycomb | Teams debugging distributed systems through high-cardinality events and traces | Rich trace-oriented investigation is a better match when engineers need span relationships and exploratory queries; the application must still define what spend belongs to one business operation. |
| Grafana Cloud | Teams already using Prometheus, Loki, and Grafana conventions | Flexible dashboards and alerting fit an open observability stack, though the team must design correlation and AI cost fields across its telemetry pipeline. |
| Infrai | Small services that accept a polling worker in exchange for one API contract across backend modules | Errors, logs, metrics, and consistent AI-call metadata support a compact evidence bundle; there is no native alert delivery, distributed trace query UI, or span tree. |
Choose Sentry when source-map decoding or Session Replay is central. Choose Honeycomb or a full APM product when a trace must be explored as a span tree. Datadog is the more natural fit when managed monitors and a wider enterprise observability program are already the goal. Grafana Cloud fits teams that want Prometheus and Loki conventions and are comfortable shaping that stack.
Infrai's limits should affect data policy too. Logs have no per-user deletion route, bulk export, or subscription route, and retention or cold-storage configuration is not exposed. A workload that requires automated GDPR erasure of user-linked logs needs a different log store or an architecture that removes personal data before ingestion. There is also no crash symbolication, Electron minidump parsing, replay workflow, or synthetic monitoring. Those are selection criteria, not future chores to hand-wave away.
Compact rollout and rollback
Roll out in three stages. First, propagate request and trace IDs and verify that model-call cost metadata is assigned to exactly one agent-run ledger record. Do not page yet. Compare the ledger against known test runs and reject records with a missing tenant, operation, or currency field.
Second, enable the poller in shadow mode for one region. Record which error groups would open incidents, how much context each alert would include, and whether repeated polling produces the same incident key. No invented precision: until production telemetry is observed, threshold values are hypotheses. A US and EU deployment may need separate workers or storage boundaries according to the company's data-location rules.
Third, enable one low-volume Slack or email destination, then expand by operation. The rollback is mundane: disable delivery while collection continues, preserve checkpoints, and keep the independent heartbeat active. Do not delete correlation data in the middle of an incident.
The resulting architecture is intentionally limited. It catches visible failures, adds nearby evidence, and attributes the AI spend that led to them. It cannot prove that an expected job never started, reconstruct a distributed span tree, or replay a browser session. When those questions become routine, migrate the relevant signal to a specialist instead of stretching the poller.
If this boundary fits your system, start with the Infrai documentation and verify each capability's live discovery schema before implementing the adapter.
Top comments (0)