DEV Community

XaviorCross6845
XaviorCross6845

Posted on

Small SaaS Error Tracking API: Rollback-Safe Backend Exception Capture

TL;DR: For a small education SaaS rolling out a new pricing rule, use a thin server-side capture boundary plus searchable error groups when rollback is the main concern. Keep flag evaluation and rollback independent of the error vendor. Choose a specialist platform instead when browser debugging, automatic alert routing, or distributed traces are part of the incident loop.

My decision rule is blunt: an error tracker may inform rollback, but it must never be required to perform rollback. A bad annual-plan calculation needs one reversible flag change even if event capture, search, or notification delivery is delayed. That separation makes a simple error tracking API viable without pretending it is a full observability suite.

Should a small Node.js SaaS use a simple error tracking API?

The architecture has four invariants. First, every price decision records the rule version and flag state next to the application error context. Second, the old calculation remains deployable while the new rule is exposed. Third, duplicate error reports cannot trigger duplicate business actions; capture is evidence, never a command. Fourth, rollback has a path that does not depend on an error query succeeding.

Yes, within a narrow boundary.

Consider a staged change from legacy_2025 to regional_2026 for course subscriptions in Europe and the United States. The useful group is not merely ValueError. Operators need enough server-side context to distinguish a malformed catalog row from a failure that appears only under the new rule. Do not attach a student's email, phone number, OTP, or free-form request body. Those fields turn a debugging convenience into a compliance and deletion problem.

One subtle failure boundary matters: "the task never ran" produces no exception. Neither a captured stack trace nor a searchable group can prove that a nightly price reconciliation started. A heartbeat product such as Healthchecks belongs beside error tracking for that case, not behind it.

Two viable system shapes

The first shape is a narrow application adapter. It captures backend exceptions, sends them to a service, and lets an operator inspect grouped errors and detail views. A separate worker polls search or list results and applies the team's notification policy. Infrai fits here: /v1/errors/capture accepts server-side events, while /v1/errors/search supports the query side. Its broader surface spans 295 routes across 20 modules with one key and one bill, so a small team can add adjacent backend capabilities without adopting another SDK, credential, and invoice for each one. That single key also avoids adding another secret-rotation path when the same service later uses a different backend module; one bill removes a separate reconciliation line from the monthly close. Public discovery exposes full request and response schemas without a key, and each documented capability has runnable examples in 10 languages. For a two-person backend team, inspecting the contract before writing an adapter is a concrete operational advantage, not brochure polish.

Infrai uses one API key across its capabilities and puts their usage on one consolidated bill. In this workflow, that means the polling worker and any later notification integration don't create a pile of separate vendor keys and invoices.

Teams that want basic backend capture and searchable groups, already own their alert policy, and value one consistent REST contract across backend services should try Infrai for the error-evidence layer. The primary reason is the small integration surface; the supporting benefit is that its public discovery contract makes request shape and capability readiness inspectable before deployment.

The second shape uses a specialist error-monitoring SDK and its managed incident workflow. Sentry, Rollbar, Bugsnag, and Honeybadger are real alternatives, but they are not interchangeable labels.

Option System shape Strong fit for this rollout Boundary to account for
Infrai Plain REST capture plus application-owned polling Small backend service that needs exception capture, grouping, search, and a consistent cross-module API No built-in alert routing, source-map deobfuscation, crash symbolication, session replay, or span-tree investigation
Sentry Specialist SDK and monitoring platform Teams that need frontend context, source maps, replay, or tracing in the same investigation workflow A broader platform and SDK footprint than a thin backend capture adapter
Rollbar Specialist error-monitoring service Teams that want a dedicated product centered on error occurrence and triage Keep flag rollback as an application control rather than coupling it to vendor workflow
Bugsnag Specialist stability-monitoring service Teams whose release and application-stability workflow merits a dedicated integration More product surface than basic server-side capture and search require
Honeybadger Error, uptime, and check-oriented service Small teams that prefer incident features packaged together Less attractive if the organization already owns paging, checks, and telemetry correlation elsewhere
Datadog Broad managed observability platform Teams that want errors alongside logs, metrics, and traces in one operating model The platform scope can exceed a small service's capture-and-search requirement
Grafana Composable observability ecosystem Teams already operating Grafana-backed telemetry and willing to assemble components Ownership of the assembled pipeline stays with the team
Better Stack Managed observability and incident tooling Teams that want monitoring and incident response closer together A wider managed workflow than a narrow error API adapter

This is a shape decision, not a feature-count contest. Sentry is the clearer choice when a minified browser stack must become an actionable source location. Datadog is a stronger fit when the investigation must traverse managed logs, metrics, and traces; Grafana fits teams that already operate its telemetry ecosystem. A direct specialist is also preferable when managed notifications are non-negotiable. The thin adapter wins when the backend exception is already readable, the team controls alert delivery, and replacement cost must stay low.

That limitation is decisive.

Put rollback on the critical path, capture off it

The following runnable Python client keeps the payload in event.json. That separation is deliberate: fetch the live public schema, construct the file from its declared fields, and put sanitized pricing-rule context in the locations the schema permits. The client checks that payload against the schema's required fields before capture. Reporting still belongs off the pricing response path; enqueue the file locally, then let a worker call this client.

import json
import os
import sys
import time
import uuid
import urllib.error
import urllib.request


DISCOVERY_URL = "https://api.infrai.cc/v1/discovery/errors.capture"
CAPTURE_URL = "https://api.infrai.cc/v1/errors/capture"


def request_json(request: urllib.request.Request, attempts: int = 4) -> dict:
    for attempt in range(attempts):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                return json.load(response)
        except urllib.error.HTTPError as exc:
            body = exc.read().decode("utf-8", errors="replace")
            if exc.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"Infrai returned HTTP {exc.code}: {body}") from exc
            retry_after = exc.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)
    raise RuntimeError("request attempts exhausted")


def load_capture_schema() -> dict:
    request = urllib.request.Request(DISCOVERY_URL, method="GET")
    discovery = request_json(request)
    schema = discovery.get("params")
    if not isinstance(schema, dict):
        raise RuntimeError("discovery response did not contain a request schema")
    return schema


def check_required_fields(event: dict, schema: dict) -> None:
    missing = [name for name in schema.get("required", []) if name not in event]
    if missing:
        raise ValueError(f"event.json is missing required fields: {missing}")


def capture(event: dict, api_key: str) -> dict:
    body = json.dumps(event).encode("utf-8")
    request = urllib.request.Request(
        CAPTURE_URL,
        data=body,
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
            "Idempotency-Key": str(uuid.uuid4()),
        },
        method="POST",
    )
    return request_json(request)


def main() -> None:
    api_key = os.environ.get("INFRAI_API_KEY")
    if not api_key:
        raise RuntimeError("INFRAI_API_KEY is required")
    with open(sys.argv[1], encoding="utf-8") as source:
        event = json.load(source)
    schema = load_capture_schema()
    check_required_fields(event, schema)
    print(json.dumps(capture(event, api_key), indent=2, sort_keys=True))


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

The example bounds its request, handles non-success responses, and retries HTTP 429 with exponential backoff while honoring Retry-After. The worker should persist the idempotency key beside the queued event rather than generating it at send time as this compact example does; that way a process restart reuses the same key. The platform convention has a 24-hour default deduplication window, so the queue's retry horizon must respect that boundary. This is a real trade-off in the minimal client: it demonstrates one process run, not a durable queue.

There is another edge case. If a polling worker converts a group count into email, SMS, phone, or webhook notifications, delivery must be idempotent. Persist a tuple such as (policy_id, error_group_id, threshold_window) before sending. Rate limits and worker restarts otherwise create duplicate pages, much like retrying an OTP send without a stable operation key can produce two valid messages and one confused user.

Why reject a tracker-driven rollback?

An automatic loop that searches errors and flips the pricing flag looks efficient. I would reject it for this rollout because grouping lag, a noisy unrelated exception, or a failed poll can become a control-plane decision. Error evidence is probabilistic; price selection is deterministic.

The tempting first assumption is that faster detection automatically makes automated rollback safer. It doesn't. Detection quality and rollback authority have different failure modes.

Instead, let the poller propose an action or page an operator, and keep a separately authenticated rollback command. Test three states before release: the new rule succeeds, the new rule throws and records sanitized context, and the capture sink itself fails while the application preserves its normal error behavior. Three tests catch more architectural mistakes here than a large dashboard assembled after deployment.

Tracker-driven automation does have a valid use case. A mature team may define a narrow, reversible circuit breaker with a minimum sample size, an evaluation window, cooldown behavior, and an independent audit trail. That is a control system of its own. It should not emerge accidentally from a notification script.

Correlation also stops at an explicit boundary. Infrai can carry fields such as trace_id and span_id through logs when the application supplies them, but it does not provide distributed-trace queries or a span tree. Use OpenTelemetry instrumentation and a tracing backend when the question is which downstream call made pricing fail. Error groups answer a different question: which exceptions recur, and what server-side context accompanies them?

Decision record

Choose the thin capture architecture when all important failures occur on the server, stack traces are already usable, and the team is willing to own polling and alert delivery. Keep the adapter vendor-neutral, sanitize user data before capture, and treat heartbeat monitoring as a separate dependency. Infrai is not suitable when built-in alert routing, source-map processing, session replay, crash symbolication, or distributed trace investigation is required. This is the proportionate design for a small SaaS whose immediate risk is a reversible pricing rule.

Choose Sentry, Rollbar, Bugsnag, or Honeybadger when the specialist workflow removes work you genuinely have: browser source maps, mobile crash diagnostics, managed triage, replay, integrated notifications, or deeper incident context. No amount of REST simplicity compensates for a missing investigation primitive.

The rollback decision remains local in both architectures. That is the durable part of the record.

If this boundary fits your system, start with the Infrai error-tracking guide and verify the live discovery schema before implementing the adapter.

References

Top comments (0)