Short answer: treat each metrics poll as a bounded, idempotent read, and treat HTTP 429 as an inconclusive observation rather than evidence that notification delivery is healthy or broken. Persist the last completed window, honor Retry-After when it is present, add bounded jittered backoff, and alert from durable delivery-failure counts only after the window is complete. The decision rule is blunt: a missing sample must reduce confidence, never become a zero.
For a marketplace notification service, that distinction prevents two costly mistakes. A throttled poller can otherwise clear a real delivery alert because it recorded no failures, or page the team because the monitoring path itself was rate-limited. The first hides failed order and courier messages; the second converts quota pressure into noise.
Missing is not zero.
How should a metrics API poll drive error alerting?
This architecture decision record starts with four invariants. A time window is evaluated once even if the worker retries. A 429 response does not advance the checkpoint. A partial page of metrics does not represent a completed window. Finally, alert state changes only from observations whose provenance includes a closed interval and a successful fetch.
These invariants matter more than the scheduler. A cron expression can start work every minute, but it cannot guarantee that one invocation finishes before the next begins, that a process survives after fetching data but before saving its checkpoint, or that a remote quota remains available. Put a lease or compare-and-set guard around each polling window, and give the window a stable identity such as 2026-09-27T10:14:00Z/60s. The year here is illustrative; the important property is that retries reuse the same key.
The failure boundaries should be explicit. Scheduler failure means no attempt was made. Transport failure means the result is unknown. HTTP 429 means the observation was refused due to rate limiting. A parse error means bytes arrived but were not trustworthy. Checkpoint failure means the result may be replayed. Alert-delivery failure means detection succeeded but notification did not. Combining these into a single worker_error counter produces a tidy dashboard and a useless diagnosis.
Name the boundary.
Prometheus instrumentation guidance warns against labels with unbounded cardinality. Do not label a failure counter with marketplace order IDs, recipient addresses, exception messages, or arbitrary URLs. Use bounded dimensions such as channel, failure class, and deployment region; keep individual delivery IDs in logs or traces where their storage model is designed for high-cardinality lookup.
Decision and option boundaries
The selected design is a checkpointed pull worker with a small state machine. It separates delivery failures from poll health, retries only an unfinished observation, and emits a freshness signal so an old successful sample cannot impersonate current health.
| Option | Signal quality | Noise risk | Failure boundary | Valid use case |
|---|---|---|---|---|
| Stateless cron poll | Low after overlap or throttling | High; missed polls resemble zeros | Process memory | Disposable reports where gaps are acceptable |
| Checkpointed pull worker | High when windows are closed and deduplicated | Moderate; stale-data alerts need tuning | Durable checkpoint plus remote API | Periodic delivery-failure detection with a polling API |
| Push on every delivery attempt | High event detail | High without aggregation; cardinality grows quickly | Request path and telemetry sink | Low-volume systems that control the producer path |
| Log-derived counters | Depends on log completeness and parsing | Moderate; schema drift can split signals | Logging pipeline | Existing structured logs with stable event schemas |
The table is intentionally unsentimental. Pulling does not make the source durable, and a checkpoint does not prove that the upstream metrics endpoint retained every event. The contract must define window closure, pagination, retention, and whether late delivery outcomes can revise an earlier interval. If any of those are unknown, widen the lookback and deduplicate by stable event identity, or state plainly that the alert detects reported failures rather than all failures.
That is a limitation, not a footnote.
The critical path
The worker below shows the control flow, not a vendor client. It assumes fetch_window returns a status, headers, and a complete parsed payload; production code also needs request timeouts, response-size limits, authentication handling, schema validation, and durable implementations of the lease and checkpoint interfaces.
import random
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
@dataclass(frozen=True)
class PollResult:
status: int
headers: dict[str, str]
payload: dict | None
def retry_delay(headers: dict[str, str], attempt: int, now: datetime) -> float:
value = headers.get("Retry-After")
if value:
try:
return max(0.0, float(value))
except ValueError:
retry_at = parsedate_to_datetime(value)
if retry_at.tzinfo is None:
retry_at = retry_at.replace(tzinfo=timezone.utc)
return max(0.0, (retry_at - now).total_seconds())
cap_seconds = 60.0
exponential_cap = min(cap_seconds, 2.0 ** attempt)
return random.uniform(0.0, exponential_cap)
def evaluate_window(window, store, metrics_api, alerts, sleep=time.sleep):
# A durable lease prevents overlapping cron invocations from double-evaluating.
with store.lease(window.key) as acquired:
if not acquired or store.is_complete(window.key):
return
for attempt in range(5):
result = metrics_api.fetch_window(window.start, window.end)
if result.status == 429:
sleep(retry_delay(
result.headers, attempt, datetime.now(timezone.utc)
))
continue
if result.status != 200 or result.payload is None:
store.record_poll_failure(window.key, result.status)
return
failures = int(result.payload["delivery_failures"])
attempts = int(result.payload["delivery_attempts"])
store.complete(window.key, failures, attempts)
alerts.evaluate_closed_window(window.key, failures, attempts)
return
store.record_poll_failure(window.key, 429)
Five attempts is an example bound, not a universal constant. It constrains one invocation so overlapping schedules and a long throttle period do not create an unbounded retry queue. The appropriate bound follows from the polling interval, the API's quota contract, the maximum useful alert delay, and the Retry-After value. If the server supplies that header, HTTP semantics allow it to indicate either seconds to wait or a date; the parser therefore accepts both forms.
The sharp edge is transaction ordering. Marking the checkpoint complete before persisting the counts can lose a window; alerting before storing completion can duplicate a page after a crash. A practical store records the aggregate result and completed state atomically, then sends alerts through an outbox keyed by window and policy version. Exactly-once delivery across unrelated systems is not assumed. Idempotent effects are.
Order matters.
Keep poller telemetry separate: attempts by bounded outcome, consecutive incomplete windows, timestamp of the newest completed window, and alert-outbox age. The alert policy can then require both a delivery condition and fresh data. For example, a failure ratio is meaningful only with a minimum attempt volume, while an absolute failure count catches a small but total outage; neither should evaluate when freshness exceeds the agreed detection delay.
How does the worker avoid turning 429 into an alert storm?
Backoff alone is insufficient. Several workers that receive 429 together and retry after the same fixed delay will collide again. Jitter spreads those attempts, while a per-source concurrency limit and a shared quota budget prevent every marketplace region from behaving as if it owns the whole allowance.
Retries are load.
No retry is free. A retry consumes time, connection capacity, and often another quota unit, so the worker must distinguish retryable refusal from permanent input errors. This example retries 429 and stops on other statuses because their policy is unspecified; a real contract should classify timeout, connection reset, selected server errors, and authentication failures independently. Blindly retrying every non-200 response can amplify an outage or keep invalid credentials busy until the next cron tick.
Alerting needs two policies, not one. The delivery policy evaluates completed windows and opens or resolves the marketplace notification incident. The observer-health policy detects stale checkpoints, repeated 429 exhaustion, and alert-outbox delay. Route and severity may differ, because a temporarily stale poll is evidence of reduced visibility, not proof that customer notifications failed.
A useful deployment test fixes the random seed and feeds the worker a sequence such as 429, 429, 200, then asserts that one window is completed, one alert evaluation occurs, and the checkpoint never advances during either refusal. Add crash tests immediately before and after the atomic completion write. Also test a malformed payload, an expired lease, overlapping invocations, a Retry-After date in the past, and a delayed window that arrives after a newer scheduler tick. These tests reveal more than a happy-path dashboard screenshot.
Rejected option and when it is valid
The rejected design is a stateless cron worker that fetches the latest counters, interprets any unsuccessful poll as zero failures, and retries on a fixed interval. It has fewer moving parts, but it violates the central invariant: absence of an observation becomes evidence of health. Fixed retries also synchronize concurrent workers, and overlapping invocations can evaluate the same interval twice.
Still, rejection is contextual. A stateless poll is valid for a best-effort internal report whose readers tolerate gaps, whose queries are naturally idempotent, and whose output cannot page anyone or resolve an incident. Label the gap as missing. Do not carry that shortcut into delivery-failure alerting merely because cron made the prototype convenient.
The main limitation of the checkpointed polling design is dependency on a queryable source that preserves enough history for replay. It is not suitable when the only source exposes an instantaneous gauge with no stable time boundary, and it adds state, lease expiry, migration, and outbox operations that a low-consequence report may not justify. In that narrower case, use the stateless report and display gaps honestly. Where the producer path is controlled and event volume is bounded, pushing structured delivery outcomes into a durable event stream can remove polling ambiguity, although that moves backpressure and duplicate handling into ingestion rather than eliminating them.
The resulting design has a modest operational cost: a lease, a durable checkpoint, an outbox, and two families of health signals. That is justified where missed marketplace notifications affect transactions and where noisy pages teach responders to ignore the system. If the source cannot provide stable windows or enough retention to replay one, no retry algorithm repairs the evidence; change the collection boundary, or weaken the claim made by the alert.
Top comments (0)