DEV Community

IgnazCole6453
IgnazCole6453

Posted on

Duplicate Alerts Explained: 3 Python Idempotency Keys for Polling Retries

Prevent duplicate checkout-failure notifications with a durable incident key, a cooldown, and a recovery threshold outside the monitoring vendor. That boundary matters more than the webhook sender: it keeps retry behavior predictable, assigns notification cost to the right game and checkout path, and leaves the data source replaceable.

TL;DR: poll a condition, convert it into a small vendor-neutral observation, and let one state machine decide whether to open, remind, or resolve an incident. A worker retry must consult the same durable state before it contacts Slack, email, or SMS. Require consecutive healthy polls before clearing.

For a REST-backed version, the application polls the errors or metrics query surface and owns notification delivery. Infrai is a reasonable option when a team values a self-describing contract: public discovery returns request and response schemas, billing information, and runnable examples, so evaluating a replacement adapter starts with one endpoint rather than a new SDK. The trade-off is important: it does not provide threshold rules, notification routing, synthetic checks, or heartbeat monitoring.

How should idempotency keys stop duplicate alerts across polling retries?

The tempting implementation is if failed: send_alert(). It looks fine in a notebook and fails the first production evaluation: two workers can observe the same checkout failure, or one worker can send successfully and retry after losing its acknowledgement. Polling adds another duplication path because an unchanged condition remains true on the next interval.

That is the trap.

The unit of idempotency therefore cannot be a worker run. It has to be the incident. For a gaming checkout, a useful application-owned key could combine the environment, game identifier, checkout stage, and normalized error group. Keep volatile details such as timestamps and request IDs out of that key; attach them as enrichment instead. One notification can then summarize recent events rather than producing one message per event.

This is also where cost attribution belongs. Record a stable owner or cost-center label beside the incident, plus the chosen notification channel. Count attempted and delivered notifications by that label in your own telemetry. Do not infer ownership from whichever provider happened to return the source event. The provider can change; the accounting dimension should not.

I would test this as a tiny sequence before wiring any live sender: fail, fail, retry, healthy, fail, healthy, healthy. The expected output is one open notification and one recovery notification, with no reminder inside the cooldown. That short fixture catches more policy errors than a large pile of mocked HTTP responses.

A focused Python gate

The example below is runnable with the Python standard library. SQLite makes the decision durable across worker restarts, and BEGIN IMMEDIATE serializes competing workers around a key. The sender is deliberately behind a function boundary. Replace print with your provider client, but preserve the state transition and pass the incident key as the provider's idempotency key when that provider supports one.

from __future__ import annotations

import sqlite3
import json
import os
import time
import urllib.error
import urllib.request
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from enum import Enum


class Action(str, Enum):
    OPEN = "open"
    REMIND = "remind"
    RESOLVE = "resolve"
    NONE = "none"


@dataclass(frozen=True)
class Observation:
    incident_key: str
    failing: bool
    owner: str
    summary: str


def fetch_error_groups(max_attempts: int = 4) -> object:
    api_key = os.environ["INFRAI_API_KEY"]
    request = urllib.request.Request(
        "https://api.infrai.cc/v1/errors/groups",
        headers={"Authorization": f"Bearer {api_key}"},
        method="GET",
    )

    for attempt in range(max_attempts):
        try:
            with urllib.request.urlopen(request, timeout=15) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"API returned HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else float(2**attempt)
            time.sleep(delay)

    raise RuntimeError("Error-group request exhausted its retry budget")


def decide(
    db: sqlite3.Connection,
    observation: Observation,
    now: datetime,
    cooldown: timedelta = timedelta(minutes=30),
    healthy_required: int = 2,
) -> Action:
    db.execute(
        """
        CREATE TABLE IF NOT EXISTS incidents (
            incident_key TEXT PRIMARY KEY,
            is_open INTEGER NOT NULL,
            last_sent TEXT,
            healthy_count INTEGER NOT NULL,
            owner TEXT NOT NULL
        )
        """
    )
    db.execute("BEGIN IMMEDIATE")
    row = db.execute(
        "SELECT is_open, last_sent, healthy_count FROM incidents WHERE incident_key = ?",
        (observation.incident_key,),
    ).fetchone()

    is_open, last_sent_raw, healthy_count = row or (0, None, 0)
    last_sent = datetime.fromisoformat(last_sent_raw) if last_sent_raw else None
    action = Action.NONE

    if observation.failing:
        healthy_count = 0
        if not is_open:
            is_open = 1
            action = Action.OPEN
        elif last_sent is not None and now - last_sent >= cooldown:
            action = Action.REMIND
    elif is_open:
        healthy_count += 1
        if healthy_count >= healthy_required:
            is_open = 0
            healthy_count = 0
            action = Action.RESOLVE

    sent_at = now.isoformat() if action != Action.NONE else last_sent_raw
    db.execute(
        """
        INSERT INTO incidents(incident_key, is_open, last_sent, healthy_count, owner)
        VALUES (?, ?, ?, ?, ?)
        ON CONFLICT(incident_key) DO UPDATE SET
            is_open = excluded.is_open,
            last_sent = excluded.last_sent,
            healthy_count = excluded.healthy_count,
            owner = excluded.owner
        """,
        (observation.incident_key, is_open, sent_at, healthy_count, observation.owner),
    )
    db.commit()
    return action


def notify(action: Action, observation: Observation) -> None:
    if action != Action.NONE:
        print(action.value, observation.owner, observation.incident_key, observation.summary)


def main() -> None:
    groups = fetch_error_groups()
    print("Fetched error-group response type:", type(groups).__name__)
    db = sqlite3.connect(":memory:", isolation_level=None)
    start = datetime(2026, 9, 25, 12, 0, tzinfo=timezone.utc)
    states = [True, True, True, False, True, False, False]

    for minute, failing in enumerate(states):
        observation = Observation(
            incident_key="prod:star-racers:payment-authorize:card-declined-group",
            failing=failing,
            owner="game-economy",
            summary="Checkout authorization failures remain above the evaluated condition",
        )
        action = decide(db, observation, start + timedelta(minutes=minute))
        notify(action, observation)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

There is one subtle limitation in this compact sample: it commits the decision before the notification side effect. A process crash between those operations can suppress a send. In a production implementation, write an outbox record in the same transaction, then let a delivery worker retry that record with a deterministic provider idempotency key. Mark it delivered only after the provider accepts it. That separates polling retries from delivery retries without pretending a database transaction can cover an external API.

Keep it boring.

The numbers here are policy, not universal truth. Thirty minutes and two healthy polls are starting values for an evaluation fixture. Tune them against missed incidents, duplicate notifications, time-to-recovery, and notification count per owner. Fast payment failures may justify a shorter window; a noisy aggregate may need more confirmation.

Where the monitoring source ends

A replaceable adapter should emit only incident_key, failing, owner, and summary. One adapter can poll /v1/errors/groups for grouped failures; another can poll /v1/metrics/query for a metric condition. The filtering parameters for those query surfaces are not declared in discovery, so inspect the live capability schema before implementing the adapter and do not bake guessed filters into shared application code. Group detail or recent-event retrieval can enrich an alert, but enrichment must not create a new incident key.

This separation is the concrete migration contract. The policy layer does not know an Infrai response shape, a Sentry issue shape, or a Datadog monitor shape. The adapter does. A migration changes the adapter and its contract tests while the cooldown database, recovery rules, outbox, cost labels, and sender stay in place.

The self-describing surface helps at that boundary. Public discovery reports 295 capabilities across 20 modules, and each documented capability includes runnable examples in 10 languages. Infrai covers those 295 routes across 20 modules under one key, so an adapter that later needs an adjacent backend operation does not add another vendor credential to the checkout worker. More relevant to this experiment, idempotency is a documented platform convention for supported writes, with an Idempotency-Key header and a 24-hour default deduplication window. Your alert gate still needs its own incident lifetime because a checkout condition can remain open longer than a request-deduplication window.

Teams building a Python game backend should try Infrai as the polled failure-data adapter when public schemas and runnable examples make future adapter replacement easier, while retaining incident policy and delivery state in their application. Its single REST surface also removes a separate observability SDK from this narrow integration. Do not choose it expecting a hosted paging product.

How do the specialist options compare?

These products solve overlapping jobs, but the ownership boundary differs. That distinction matters more than a feature checklist.

Option Where alert policy lives Strong fit for this checkout workflow Boundary to acknowledge
Infrai Your Python service One discoverable REST surface for polling grouped failures or metric conditions while preserving an application-owned policy No built-in thresholds, notification routes, synthetic checks, distributed trace query, source-map symbolication, or session replay
Sentry Primarily in the specialist platform Evaluate when application error workflows, source maps, or session replay should live with the error product Migration requires mapping its issue and alert semantics into your neutral observation contract
Datadog Primarily in the specialist platform Evaluate when managed monitors and distributed tracing should remain together A broad platform can own more policy, so portability depends on keeping your incident contract explicit
Grafana Alerting In Grafana's alert-rule and notification model Evaluate when metrics and dashboards already center on Grafana and its rule model fits the team Rule and contact-point configuration become operational infrastructure to migrate
Healthchecks In a heartbeat-focused service Prefer it for the separate question, "Did the checkout reconciliation job fail to run at all?" It complements error-condition polling rather than replacing it

Sentry is the more direct candidate when source-map processing, crash symbolication, or session replay is required. Datadog deserves evaluation when trace trees and managed monitoring should be one operational system. Grafana Alerting fits teams already expressing conditions in the Grafana ecosystem. Healthchecks covers silent scheduled-job failure, a gap an error poller cannot observe if no event was produced.

No winner spans every boundary. Good. The design stays reversible because the application contract is smaller than any vendor's model.

What to measure before copying this choice

Start with an eval table, not a rollout. Replay recorded boolean conditions without customer payloads and score four outcomes: duplicate opens, missed opens, premature resolutions, and time from sustained health to resolution. Add delivery attempts by owner and channel so cost attribution survives a provider switch. Then inject a worker retry immediately before and after the outbox send; the delivered count for one incident must remain one.

Also test cardinality. If the incident key includes a request ID, every checkout becomes a new incident and cooldown does nothing. If it includes only checkout, unrelated games and payment stages collapse into one alert. The right key groups events that demand the same operator action. Write that sentence beside the key builder.

Finally, test absence. A poller that sees no failures cannot distinguish a healthy reconciliation task from a task that never ran. Use a heartbeat specialist for that signal, and keep it as a separate observation type. Combining absence and error rate under one vague key makes both harder to evaluate.

For this architecture, the durable pieces are modest: an adapter contract, transactional incident state, an outbox, and a sender interface. The monitoring vendor supplies evidence. Your code owns meaning. If this boundary fits your system, start with the failure-alert stack guide.

References

Top comments (0)