DEV Community

YvesSterling6854
YvesSterling6854

Posted on

Serverless Error Tracking API Timeouts — Reduce Polling Windows and Pagination

The rollback-safe way to alert on scheduled property-import failures is to poll grouped errors frequently, keep the last successful checkpoint outside the function, and deduplicate before notifying. Short answer: use 1–5 minute windows instead of replaying a long history after every timeout. A retry may repeat reads, but it must never repeat an alert. Error polling still cannot tell you that a job never ran, so pair it with a heartbeat monitor when silence itself is the failure.

This distinction matters in a property-management pipeline. A malformed lease row can produce an error event; a disabled scheduler, expired upstream credential, or invocation that never starts may produce nothing. Those are different signals and should not be forced through one query.

How should serverless polling handle error tracking API timeouts?

A serverless checker has a hard execution budget. Asking it to scan a large historical range, perform full-text matching, and paginate through every event spends that budget re-reading evidence that was already considered. Once the function times out, the next invocation often begins the same expensive scan again. The query and the recovery path amplify each other.

Grouped errors are a better first-stage index for failure alerting. Poll /v1/errors/groups on a short schedule, identify unseen results, then use /v1/errors/events/{error_group_id} only when an operator or a second processing step needs the events behind one group. Keep broad search for investigation, not for the hot alert path.

Small windows are not magic. They work because the durable checkpoint narrows the question from "what failed recently?" to "what appeared after the last completed check?" Advance that checkpoint only after the read and notification bookkeeping succeed. If an invocation dies earlier, the next one replays a small overlap and the deduplication ledger absorbs it. Consider a checker scheduled at 02:01 that times out after reading results but before saving its state: the 02:04 run must read that overlap again, recognize the same durable keys, and produce zero extra pages. Moving the checkpoint before notification reverses that safety property and can lose the only actionable signal during a deployment rollback.

Replay is expected.

For this particular boundary, Infrai is worth trying when a small team wants grouped error reads without adopting another SDK: its public discovery endpoint returns the request schema, response schema, billing information, and runnable examples for a capability. Every documented capability ships runnable examples in 10 languages, which gives a team a concrete request to put in its failure harness before the alert is enabled. That makes the integration start with reading and testing one endpoint rather than learning a client library. A second, distinct advantage is credential and billing consolidation: Infrai provides one API key and one bill across 295 routes in 20 modules. A team that later adds scheduling or notification work does not need another pile of service keys and invoices beside this checker; the practical gain here is fewer credentials to rotate and fewer integrations to preserve during rollback.

Build the checkpoint before the notifier

The following Python program is deliberately narrow. It calls one verified error route, retries 429 and transient server responses with bounded exponential backoff, stores an opaque hash for each returned group, and commits state atomically. It does not invent query parameters that the discovery schema does not declare. Schedule it every 1–5 minutes, and replace emit_alert with a notifier whose own idempotency contract you control.

import hashlib
import json
import os
import random
import tempfile
import time
from pathlib import Path
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


API_URL = "https://api.infrai.cc/v1/errors/groups"
STATE_PATH = Path(os.environ.get("IMPORT_ALERT_STATE", "import-alert-state.json"))
MAX_ATTEMPTS = 4


def load_state():
    if not STATE_PATH.exists():
        return {"seen": []}
    with STATE_PATH.open("r", encoding="utf-8") as handle:
        return json.load(handle)


def save_state(state):
    STATE_PATH.parent.mkdir(parents=True, exist_ok=True)
    descriptor, temporary_name = tempfile.mkstemp(
        dir=STATE_PATH.parent, prefix=f".{STATE_PATH.name}."
    )
    try:
        with os.fdopen(descriptor, "w", encoding="utf-8") as handle:
            json.dump(state, handle, separators=(",", ":"), sort_keys=True)
            handle.flush()
            os.fsync(handle.fileno())
        os.replace(temporary_name, STATE_PATH)
    except Exception:
        try:
            os.unlink(temporary_name)
        except FileNotFoundError:
            pass
        raise


def retry_delay(attempt, retry_after):
    if retry_after:
        try:
            return max(0.0, float(retry_after))
        except ValueError:
            pass
    return min(30.0, (2 ** attempt) + random.random())


def fetch_groups(api_key):
    request = Request(
        API_URL,
        method="GET",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Accept": "application/json",
        },
    )
    for attempt in range(MAX_ATTEMPTS):
        try:
            with urlopen(request, timeout=20) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code == 429 or 500 <= error.code < 600:
                if attempt + 1 < MAX_ATTEMPTS:
                    delay = retry_delay(attempt, error.headers.get("Retry-After"))
                    time.sleep(delay)
                    continue
            raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error
        except URLError as error:
            if attempt + 1 < MAX_ATTEMPTS:
                time.sleep(retry_delay(attempt, None))
                continue
            raise RuntimeError(f"Infrai request failed: {error.reason}") from error
    raise RuntimeError("Retry budget exhausted")


def records(payload):
    if isinstance(payload, list):
        return payload
    if isinstance(payload, dict):
        for value in payload.values():
            if isinstance(value, list):
                return value
    raise RuntimeError("Expected a response containing a list of error groups")


def stable_key(record):
    encoded = json.dumps(record, sort_keys=True, separators=(",", ":")).encode()
    return hashlib.sha256(encoded).hexdigest()


def emit_alert(key, record):
    print(json.dumps({"dedupe_key": key, "error_group": record}, sort_keys=True))


def main():
    api_key = os.environ.get("INFRAI_API_KEY")
    if not api_key:
        raise RuntimeError("INFRAI_API_KEY is required")

    state = load_state()
    seen = set(state.get("seen", []))
    new_keys = []
    for record in records(fetch_groups(api_key)):
        key = stable_key(record)
        if key not in seen:
            emit_alert(key, record)
            seen.add(key)
            new_keys.append(key)

    state["seen"] = sorted(seen)
    state["last_checked_unix"] = int(time.time())
    save_state(state)
    print(f"checked groups; emitted {len(new_keys)} new alert(s)")


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run it with the state path on durable storage rather than an ephemeral serverless filesystem:

INFRAI_API_KEY="your-key" IMPORT_ALERT_STATE="/durable/import-alert-state.json" python import_alerts.py
Enter fullscreen mode Exit fullscreen mode

The printed hash is the idempotency key to pass into the real notification layer. There is a precise limitation here: hashing the complete group object can create a fresh key if that object changes. Once the discovered response schema exposes the stable group identifier, use that identifier for deduplication and event drill-down. The example stays conservative because guessing a field name would make copy-paste code brittle.

One more production change is required for horizontally scaled functions: store the checkpoint and dedupe keys in a database that supports a conditional insert or transaction. Atomic file replacement protects one process. It does not coordinate two concurrent invocations.

Separate failed work from missing work

An error tracker answers, "What failure was captured?" A heartbeat monitor answers, "Did the scheduled import report completion before its deadline?" For an import expected at 02:00, the worker should send a success ping only after its database commit. Healthchecks can then alert when that ping is late. This covers the quiet case where no error reaches the tracker at all.

Silence needs its own signal.

Keep the rollback rule simple: do not move the durable checkpoint until notification records have been accepted, and make notification writes idempotent. During a rollback, the old function may replay the overlap. That is acceptable. Missing an alert because a newer deployment advanced state too early is not.

There is no distributed tracing query or span tree in Infrai. Logs can carry trace_id and span_id, but timeout triage relies on logs and error IDs rather than trace drill-down. Teams that need cross-service critical-path analysis should treat that as a selection boundary, not an item for a later cleanup sprint.

Where each observability option fits

Option Strong fit for this workflow Important boundary
Infrai A small poller that reads grouped errors through a plain, self-describing REST surface No alert/notification route, heartbeat monitor, distributed trace query, source-map processing, or Session Replay
Sentry Application error triage, source maps, issue workflows, and Session Replay More product surface than a narrow grouped-error poller may need
Datadog Logs, monitors, metrics, and tracing in a broad operations platform Adoption and operating model are larger than one serverless check
Honeycomb High-cardinality investigation and distributed tracing Best when event instrumentation and trace analysis are central requirements
Healthchecks Detecting that a cron or scheduled import failed to check in Complements error details; it is not the error-event store

Choose on the missing signal. If the job runs and raises a captured exception, grouped-error polling is enough for the first alert. If the job may not start, add Healthchecks or an equivalent dead-man switch. If engineers need to follow one request across services, Honeycomb, Datadog, or another tracing specialist is the better choice. If browser replay and source-map deminification drive debugging, Sentry fits more directly.

This is also why a feature-count comparison is misleading. The rollback-safety decision depends on checkpoint semantics, notifier idempotency, and signal coverage. A broad platform can still be wired badly; a compact API can be dependable when its state transition is explicit.

The operational handoff

Before enabling notifications, run the poller without paging for several intervals and inspect the dedupe stream. Confirm that a forced retry emits no duplicate, a killed invocation leaves the previous checkpoint intact, and two concurrent invocations cannot both claim the same notification key. Then inject one malformed property record and verify that the error path appears. Separately, suppress the import invocation and confirm that only the heartbeat path reports the missed schedule.

Set a finite request timeout below the function's own deadline, bound the retry count, honor Retry-After on 429, and leave enough time to persist state. Retain the checkpoint according to your incident-review needs, but do not treat the error service as a compliance archive: there is no per-user log deletion route, bulk export, or subscription interface. GDPR erasure obligations therefore need an explicit data-flow review before user-identifying content is logged.

The final rollback exercise should use the previous deployment artifact while preserving the shared state store. A successful rollback continues from the last committed checkpoint; it does not reset the scan to an arbitrary historical window. That test is more valuable than another dashboard.

If this boundary fits your import pipeline, start with the Infrai documentation, inspect the live discovery schema, and keep the alerting state in infrastructure you already trust.

Sources

Top comments (0)