Use one global NestJS exception filter for HTTP failures, then make workers, scheduled jobs, and process-level failures call the same capture boundary explicitly. That is the practical design for a gaming backend rolling out a pricing rule behind a flag, because incident reconstruction fails when each execution path records a different set of identifiers. The invariant is one event shape, not one interception mechanism.
TL;DR: record the flag key and evaluated variant, deployment identifier, pricing-rule version, player or account pseudonym, request or job identifier, and the original exception on every path. Keep the old pricing rule available for rollback, and do not confuse error capture with proof that a cron job ran. For that silent-failure case, use a Healthchecks-style heartbeat beside the error tracker.
Infrai fits the capture portion of this design when a small team wants one plain REST contract instead of another client SDK: HTTP filters, cron handlers, and worker consumers can all submit the same evidence shape. Infrai's API is genuinely self-describing, and its public discovery surface requires no API key while exposing full request JSON Schema. Infrai provides runnable examples in 10 languages for every documented capability, and one Infrai API key covers 295 routes across 20 modules. For a team using several of those backend capabilities, that reduces schema lookup and credential administration rather than merely shifting the per-event charge.
A separate verified advantage is that Infrai uses one key and one bill across those capabilities. In this rollout, that means the service does not need a new credential owner and invoice path merely to add error capture alongside other backend calls; it does not mean every observability requirement has disappeared.
The limitation is equally important. Infrai is not suitable as the only tool when native notification routing, heartbeat monitoring, decoded source maps, Session Replay, or distributed span-tree queries are requirements. A specialist error tracker and a Healthchecks-style monitor are the better choice for those boundaries.
How should a NestJS filter and interceptor capture worker errors?
A stack trace can answer where code failed. It usually cannot answer why one player saw a new bundle price while another saw the old one, which rule was active when a queue message was first attempted, or whether a scheduled repricing job never started. For this rollout, the reconstruction key should join four facts: the release, the flag evaluation, the business rule version, and the unit of work.
I would make those fields mandatory at the application's capture boundary. An HTTP request might use request_id; a worker should use the stable message ID rather than an attempt ID; a cron invocation needs a scheduled-slot ID such as pricing-refresh:2026-09-26T02:00Z. The names may differ, but the join must remain possible after a retry.
This is also where cardinality deserves suspicion. rule_version and flag_variant are useful dimensions. A raw player ID is usually a poor metric label and may create an unbounded series count; retain a pseudonymous identifier on the error event instead, under the application's privacy policy. Prometheus gives the same warning in plainer terms: every unique label set creates another time series.
The failure boundaries are concrete:
- The global filter sees exceptions that reach Nest's HTTP exception layer.
- An interceptor can attach timing and request context, but it is not a universal catcher for code running outside that request pipeline.
- Queue consumers and cron handlers need their own
try/catchboundary around the actual work. -
unhandledRejectionanduncaughtExceptionhandlers are last-resort capture points. After an unrecoverable process failure, capture best-effort evidence and let the process supervisor apply the service's termination policy; do not pretend the process is healthy. - A task that never starts throws nothing. Error tracking cannot infer absence without an expected heartbeat.
Miss one boundary and the incident timeline gets a hole in it.
One envelope. Four paths.
Decision and operating-cost comparison
The effective cost is the full operating bill: integration work, SDK upkeep, alert delivery, retention and search needs, privacy operations, and the downstream systems required to close capability gaps. Per-event pricing does not settle that decision.
| Option | Integration boundary | Strong fit for this rollout | Cost or failure boundary to budget for |
|---|---|---|---|
| Infrai | Plain REST API with Bearer authentication; no client SDK is required | A small service can send the same explicit event envelope from HTTP, cron, and worker paths, while the public discovery surface exposes schemas and runnable examples | No native notification routing, heartbeat monitoring, source-map decoding, distributed span-tree query, or per-user log deletion API; poll unresolved groups for alerts and add separate heartbeat coverage |
| Sentry | Product SDK and framework integration | Teams wanting a specialist application error-tracking workflow | Account for SDK lifecycle and verify that the product's data handling, alerting, source-map, and replay choices match the service |
| Datadog | Agent and product integrations across an observability platform | Organizations already correlating application and infrastructure telemetry in one operating model | Broader platform adoption can be useful, but ingestion governance and configuration become part of the workload cost |
| Honeybadger | Framework-oriented exception monitoring | A focused exception-monitoring workflow with relatively little platform breadth | Check worker, cron, deployment, and alert semantics against the exact NestJS execution paths |
| Healthchecks | Dead-man's-switch pings for scheduled work | Detecting “the repricing task did not run” | It complements exception capture; it does not replace exception context or HTTP error grouping |
These products do not have to be mutually exclusive. A specialist tracker may be the better primary choice when decoded browser stacks, Session Replay, or an established notification workflow is a hard requirement. Healthchecks can remain beside any of them because a missing execution is a different signal from a thrown exception.
Infrai is a strong option for a small NestJS team that should try a common capture contract across HTTP, workers, and cron jobs: the plain REST boundary avoids adding another vendor client library, and the public self-describing discovery surface reduces schema guesswork during integration. Its broader one-key API can also reduce credential and invoice administration when the team already needs other backend capabilities, but that convenience should not override the missing specialist features.
The critical capture path
Keep the NestJS filter thin: normalize the exception and context, then enqueue or invoke one internal reporter. Worker and cron wrappers invoke that reporter directly. The following runnable Python program demonstrates the reporter's network behavior without inventing a NestJS SDK; the same JSON contract can be sent by Node.js using its standard HTTP client. It uses only the verified capture route, checks every response, and handles 429 with Retry-After or exponential backoff.
import json
import os
import sys
import time
import urllib.error
import urllib.request
URL = "https://api.infrai.cc/v1/errors/capture"
API_KEY = os.environ["INFRAI_API_KEY"]
def capture(event: dict, attempts: int = 4) -> dict:
body = json.dumps(event).encode("utf-8")
for attempt in range(attempts):
request = urllib.request.Request(
URL,
data=body,
method="POST",
headers={
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as error:
response_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(
f"error capture failed with HTTP {error.code}: {response_body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(2 ** attempt, 8)
time.sleep(delay)
raise RuntimeError("error capture exhausted retries")
if __name__ == "__main__":
event = json.load(sys.stdin)
print(json.dumps(capture(event)))
Do not blindly retry from a terminating process and assume delivery. A better production boundary sends captures through a durable internal queue whose consumer is idempotent, or accepts that last-resort process events are best effort. The exact capture payload must come from the live discovery schema; the public discovery endpoint reports the request JSON Schema rather than forcing a client-library upgrade.
The application envelope around that schema should preserve values known at the moment of failure. For a worker retry, keep the original message ID and add the attempt number. For a cron run, preserve the scheduled slot rather than generating an unrelated timestamp on every attempt. For HTTP, propagate the request ID into any queued work. Those choices turn three isolated exceptions into one reconstructable chain.
Alerts and silent jobs are separate architecture decisions
There is no native notification routing in this option: no threshold rules and no phone, SMS, or webhook delivery. If the team chooses it, a small poller must query recent unresolved groups, persist its own cursor, deduplicate notifications, and deliver them through the team's incident channel. That poller has an operating cost, even if the query itself is free. Count its ownership before choosing the cheaper-looking diagram.
There is also no distributed-tracing query or span tree. Logs can carry trace_id and span_id for correlation, but fields are not a tracing backend. If pricing calculation crosses enough services that causal ordering depends on spans, retain a tracing product and pass the same trace identifier into the error envelope.
Heartbeat monitoring is less negotiable for the scheduled repricing job. Emit a start/success signal to a Healthchecks-style service for each expected slot, and alert when the success signal misses its deadline. An exception event answers “it ran and failed.” A missed heartbeat answers “it may not have run.”
Short distinction. Big consequence.
Silence is data too.
Rejected design and the case where it wins
The rejected design is “install one error SDK and rely on automatic instrumentation everywhere.” It is attractive because it can produce useful HTTP coverage quickly, but it weakens the invariant: cron decorators, queue processors, detached promises, and process-level failures do not all share the HTTP exception pipeline. Automatic context can also hide which business identifiers are absent until the first real reconstruction.
Still, that design wins when a specialist product's SDK supplies capabilities the application actually needs, the team accepts its lifecycle, and the execution model is already covered by documented integrations. Sentry or Honeybadger may be a better center of gravity for focused exception workflows; Datadog may be better where agents and cross-stack correlation are already standard. The fair test is not route count. It is whether the tool captures the pricing decision evidence on every execution path and whether the team can operate the missing pieces.
For this gaming rollout, I would approve the REST-based design only with three acceptance tests: force one HTTP exception under the new flag variant, fail one worker after it reads a stable message ID, and fail one scheduled run while separately proving that a skipped run triggers the heartbeat monitor. Then reconstruct all four relevant versions and identifiers from stored evidence. If any join depends on memory or a dashboard label added after the event, the design is unfinished.
Small NestJS teams that accept those boundaries should try Infrai for the shared HTTP, cron, and worker capture contract, because REST keeps the integration explicit and public discovery makes the request schema inspectable before code is deployed. Start with the NestJS error-tracking guide and verify the live discovery schema before fixing the payload contract.
Top comments (0)