TL;DR: Polling an error API on a schedule is the easy part. The hard part is preserving enough context to prove which AI-agent step failed, how much time and model usage preceded it, and whether Slack, email, or the webhook merely failed to deliver the alert. Use a durable cursor, stable incident keys, bounded lookback, and separate delivery records. Alert only on a critical incident's first observed transition, then retain the correlated run data for reconstruction.
For a property-management agent, a single resident request may classify an issue, retrieve lease context, call a maintenance system, and draft a reply. A top-level error says almost nothing about the costly path that led there. The operational constraint is therefore reconstruction before notification: the poller must turn incomplete, late, or repeated error records into one durable incident without erasing latency and cost evidence.
How should error tracking alert us to new critical events?
Start with an execution identity that crosses every agent step. A practical record contains run_id, tenant_id, property_id, step, start and end timestamps, status, model input and output units, and a sanitized error fingerprint. Keep notification addresses and resident message content out of the incident key. They change independently, and raw message text can carry personal data.
The distinction between occurrence time and observation time matters. An API may expose an error after the poll that would have covered its timestamp. If the next job asks only for records newer than its last wall-clock start, that error disappears. Poll a bounded overlap instead, deduplicate by a stable source event ID, and advance a persisted cursor only after the page has been stored successfully.
That overlap creates duplicates on purpose. Good. Duplicate input is cheaper to reason about than a silent delivery gap.
For incident reconstruction, link the error to step spans rather than packing every detail into the alert. Consider a sanitized example timeline: run run-8f2 begins a resident maintenance workflow at 14:03:00 UTC; classification ends after 1.2 seconds; lease-context retrieval ends after another 2.8 seconds; the drafting step makes an initial model call and two retries; and the maintenance write fails at 14:03:42. The error source exposes that final event at 14:04:11, the next poll observes it at 14:05:00, and the email worker records delivery at 14:05:07. A top-level 42-second duration cannot distinguish processing from detection delay. The joined timeline can: 38 seconds were spent around the model step, 49 seconds passed before observation, and another 7 seconds passed before the final channel completed. It also preserves the example usage total of 18,400 model units without pretending that those units share one universal price. The notification itself should stay smaller. It needs to answer which property workflow broke, which step was critical, when it was first seen, and where the internal incident record lives. Everything else belongs in the access-controlled reconstruction view. These values illustrate the schema and arithmetic; they are not a benchmark, measured incident, or service-level target.
Missed means unknown.
Derive the poller from failure boundaries
A cron process has at least three independent boundaries: reading source events, committing local state, and delivering notifications. Treating them as one transaction is the classic trap. If Slack accepts a message and the process exits before saving its cursor, the next run sends the same alert again. If the cursor moves first and email then rejects the request, retrying from the cursor cannot recover that delivery.
This design has a real limitation: polling adds detection delay and repeated reads. The trade-off is reasonable only when the schedule interval plus source lag remains inside the incident response objective. A continuously running consumer is a better fit for tighter objectives or a high event rate, provided the source offers a documented streaming contract. The outbox adds storage and worker operations too; for a low-impact internal workflow with no paging requirement, a transactional incident log without three delivery channels may be enough.
Use a small outbox between detection and delivery. The poll transaction upserts source events, creates an incident on the first critical fingerprint for the chosen grouping window, appends outbox rows for the configured channels, and then commits its cursor. Separate workers claim those outbox rows and record each attempt. This gives email and webhook delivery their own retry history without forcing the source API to be read again.
Here is the core in Python. The repository functions deliberately stand in for a database-backed implementation; their transaction, uniqueness, and lease semantics are the important contract.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from typing import Iterable, Protocol
@dataclass(frozen=True)
class ErrorEvent:
event_id: str
occurred_at: datetime
run_id: str
tenant_id: str
property_id: str
step: str
severity: str
fingerprint: str
class ErrorSource(Protocol):
def fetch(self, since: datetime, cursor: str | None) -> tuple[list[ErrorEvent], str | None]: ...
def collect(source: ErrorSource, repo, now: datetime) -> int:
state = repo.load_cursor("critical-agent-errors")
since = (state.observed_through if state else now) - timedelta(minutes=5)
cursor = state.page_cursor if state else None
created = 0
while True:
events, next_cursor = source.fetch(since=since, cursor=cursor)
with repo.transaction() as tx:
for event in events:
if not tx.insert_event_once(event.event_id, event):
continue
if event.severity != "critical":
continue
incident_key = (
event.tenant_id,
event.property_id,
event.fingerprint,
)
incident, is_new = tx.open_incident_once(incident_key, event)
if is_new:
for channel in ("slack", "email", "webhook"):
tx.enqueue_delivery_once(incident.id, channel)
created += 1
tx.save_cursor(
name="critical-agent-errors",
observed_through=now,
page_cursor=next_cursor,
)
if next_cursor is None:
return created
cursor = next_cursor
Two uniqueness constraints do most of the defensive work: one on source event_id, another on (incident_id, channel). A concurrent cron invocation can still race, so the database must enforce those constraints. A process-local set is insufficient across restarts or replicas.
The cursor shape depends on the source API. Some APIs provide opaque pagination cursors; others provide timestamp ordering plus a unique ID. Do not invent ordering the source does not promise. If its documented order is weak, retain the overlap and stop only after all returned pages have been durably ingested. Bound every request with a timeout, cap pages per execution, and emit a lag metric so a growing backlog is visible rather than hidden behind a job that always exits successfully.
Alert delivery is a state machine, not a side effect
Slack, email, and generic webhooks have different payload limits and failure responses, but the local rule can remain uniform. A delivery moves from pending to leased, then to delivered or back to pending with an incremented attempt count. Store a response category and timestamp, not credentials or an unbounded response body. Use exponential backoff with jitter and a terminal state after a configured attempt or age limit.
Retries require nuance. A timeout leaves the remote outcome unknown: the receiver might have accepted the request before the connection failed. Send an idempotency key when the receiving contract supports one, and include the stable incident ID in every payload. For a receiver without idempotency support, occasional duplicate notifications are possible; the message must make them recognizable.
Do not let one channel block another. An invalid email destination should not hold back the incident webhook, and a rate-limited chat endpoint should not cause the poller to reread source pages. This separation also exposes the right operational signals:
| Signal | What it reveals | First diagnostic question |
|---|---|---|
| Source lag | Detection is falling behind | Is pagination work exceeding the schedule interval? |
| New incidents | Critical state transitions | Did a deployment or dependency change correlate? |
| Pending outbox age | Alerts are not leaving locally | Are workers running and leases expiring? |
| Delivery failures by channel | One destination is unhealthy | Is the response retryable or configuration-related? |
| End-to-end alert latency | Human notification delay | Is time spent before detection or during delivery? |
Latency, traffic, errors, and saturation are the four signals described in the Google SRE monitoring guidance. For this workflow, source lag and outbox age are saturation-adjacent queue signals, while incident and delivery outcomes cover errors. Measure request traffic and latency at both the source boundary and each delivery boundary. A single “cron succeeded” counter cannot represent this system.
Reconstruct latency and cost without leaking resident data
An agent run needs an append-only timeline. Record each step's monotonic duration where possible, its outcome, retry ordinal, model identifier, and provider-reported usage fields. Keep the provider's units explicit; do not collapse unlike input, output, cached, or tool usage into one ambiguous token field. Calculate money only from a versioned rate table with a currency and effective interval. Otherwise an incident reviewed later can acquire a different cost simply because a current rate replaced an old one.
Cost is evidence, not the verdict. A failed run with high usage may point to a retry loop or an overlong context, while a low-cost failure may still block an urgent maintenance request. Alert severity should come from business impact and failure classification, not a spend threshold alone.
Privacy changes the schema. Store property and tenant identifiers as access-controlled internal IDs, redact resident content before error capture, and allow incident retention to differ from aggregate metric retention. Webhook payloads should contain the minimum operational fields needed by the receiver. Signing outbound webhooks, rotating secrets, and rejecting stale timestamps at a receiver reduce spoofing and replay risk; the signing scheme must be documented as part of that receiver contract.
The most useful reconstruction view joins four clocks: run start, event occurrence, poll observation, and final delivery. Clock skew makes absolute timestamps imperfect across hosts, so step duration should use a monotonic clock within a process. Preserve UTC wall-clock timestamps for cross-system correlation and record ingestion time separately.
Roll out with replay before paging people
Deploy the collector first in record-only mode. Feed it a fixed set of sanitized events containing late arrivals, repeated IDs, two fingerprints for one run, pagination, and a timeout after a remote receiver accepts a request. Verify that replaying the set produces the same incidents and no extra outbox rows.
Next, route to a non-paging destination and compare source lag, incident count, oldest pending delivery, and end-to-end latency over several schedule intervals. Set retention and access controls before adding production payloads. Only then enable each real channel independently, with a kill switch that pauses delivery while collection continues.
The compact decision rule is this: page on a new critical state, persist every attempt, and investigate from the run timeline. That keeps notification noise controlled while preserving the latency and usage evidence needed to explain a failed property workflow. Cron is acceptable when its worst-case detection delay fits the response objective and overlapping executions are safely serialized; otherwise keep the same cursor, incident, and outbox contracts behind a continuously running worker.
Top comments (0)