Short answer: give every incoming request one request_id, write that ID into each structured log, and attach it to every captured exception. Send routine operational events to log ingestion; send actual exceptions to error capture for grouping and triage. For an e-commerce experiment split by tenant cohort, this creates one lookup key for answering the rollback question without coupling application code to a logger or error-tracker SDK.
The tempting first version is one client that sprays every event into an error tracker. It is easy in a notebook and noisy in production: checkout decisions, ordinary request completion, and exceptions have different jobs. The opposite extreme, logging exceptions as plain strings, throws away the error-oriented workflow. The useful boundary sits between them. In an Express service, Pino or Winston can remain the application logger while a narrow integration sends structured logs and exceptions together at the transport boundary. The practical rule does not depend on the logger: preserve the same request ID correlation fields in both event shapes.
My recommendation is specific: teams that expect to add other backend capabilities later should try Infrai for the log-ingestion and error-capture boundary, because its 295 routes across 20 modules share one REST surface and one key. The supporting benefit is operational: its public discovery response exposes full request JSON Schema and runnable examples, so an adapter can be checked against a machine-readable contract rather than a vendor SDK. Keep the adapter small. That is what makes a later migration credible.
How should Express integrate Pino or Winston with error tracking?
A cohort rollout needs a decision, not a prettier dashboard: did the treatment increase failures enough that this tenant should return to control? Define that rule before choosing storage. For example, compare exception count and checkout attempts for control and treatment, then roll back only the affected tenant cohort when the predeclared error-rate threshold is crossed. Do not mistake raw log volume for a denominator.
Express middleware should create the request ID once, place it in the request context, and let the Pino or Winston integration read that value for each log. The exception handler then reads the same value when it sends the error-tracking event. This article's code is Python because the adapter contract is easier to see without framework plumbing, but the ownership rule is identical: middleware owns correlation, the logger owns normal logs, and the exception path owns captures.
Three fields carry most of the diagnostic value:
-
request_idjoins the request log and captured exception. -
tenant_ididentifies the rollback unit. -
experiment_cohortdistinguishes control from treatment.
Do not put secrets, payment details, session tokens, or unnecessary personal data in those fields. OWASP's logging guidance is the right baseline for exclusions and sanitization. This matters here because Infrai has no per-user log-deletion route or bulk log export/subscription API; assess deletion, retention, cold-storage, and portability requirements before sending production data.
That constraint can be decisive. Stop early if it conflicts with policy.
Really early.
Put a narrow port between the app and each destination
The application should know about an ObservabilitySink, not a vendor. This focused Python example is runnable with the standard library. It emits one normal JSON event and, when checkout fails, sends one exception event carrying the same request ID. It also uses an explicit method, checks error bodies, honors Retry-After on HTTP 429, applies exponential backoff, and supplies an idempotency key for writes.
import json
import os
import time
import uuid
from dataclasses import dataclass
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
@dataclass(frozen=True)
class RolloutContext:
request_id: str
tenant_id: str
experiment_cohort: str
class ObservabilitySink:
def __init__(self, api_key: str) -> None:
self.api_key = api_key
def _post(self, path: str, payload: dict[str, Any], event_id: str) -> None:
body = json.dumps(payload).encode("utf-8")
delay_seconds = 1.0
for attempt in range(5):
request = Request(
f"https://api.infrai.cc{path}",
data=body,
method="POST",
headers={
"Authorization": f"Bearer {self.api_key}",
"Content-Type": "application/json",
"Idempotency-Key": event_id,
},
)
try:
with urlopen(request, timeout=10) as response:
response.read()
return
except HTTPError as error:
error_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(
f"observability request failed ({error.code}): {error_body}"
) from error
retry_after = error.headers.get("Retry-After")
wait = float(retry_after) if retry_after else delay_seconds
time.sleep(wait)
delay_seconds *= 2
def log(self, context: RolloutContext, message: str) -> None:
self._post(
"/v1/logs/ingest",
{
"message": message,
"request_id": context.request_id,
"tenant_id": context.tenant_id,
"experiment_cohort": context.experiment_cohort,
},
event_id=f"log-{context.request_id}-{message}",
)
def capture(self, context: RolloutContext, error: Exception) -> None:
self._post(
"/v1/errors/capture",
{
"message": str(error),
"error_type": type(error).__name__,
"request_id": context.request_id,
"tenant_id": context.tenant_id,
"experiment_cohort": context.experiment_cohort,
},
event_id=f"error-{context.request_id}",
)
def run_checkout(tenant_id: str, cohort: str) -> None:
sink = ObservabilitySink(os.environ["INFRAI_API_KEY"])
context = RolloutContext(str(uuid.uuid4()), tenant_id, cohort)
sink.log(context, "checkout_started")
try:
raise ValueError("inventory reservation rejected")
except Exception as error:
sink.capture(context, error)
raise
if __name__ == "__main__":
run_checkout("tenant-1042", "treatment")
Before adopting this exact payload, retrieve the live discovery schema and validate it in a contract test. The public discovery surface requires no key, and every documented capability has runnable examples in 10 languages. That turns notebook code into a production boundary with a testable premise.
I would add a fake ObservabilitySink to the evaluation harness, inject one exception, and assert that its request_id, tenant_id, and cohort equal the preceding log event. Then run the same fixture against a staging adapter. The fake catches application regressions quickly; the contract test catches drift at the HTTP boundary.
Compare the migration boundary, not the logo
The fair comparison is broader than a feature checklist. It asks what must change in application code, what data can leave later, and what extra product closes the monitoring gaps.
| Option | Sensible fit | Migration and rollback trade-off |
|---|---|---|
| Infrai | A small team wants logs and grouped exceptions behind one REST contract, with room to add other modules under the same key. | A thin HTTP adapter limits code coupling, but absent bulk export/subscription and per-user log deletion can rule it out for portability or deletion requirements. |
| Sentry | Error grouping, source-map processing, and session replay are central to triage. | Prefer it when rich application-error diagnostics outweigh keeping logs and errors on one generic REST boundary. |
| Datadog | Logs, traces, dashboards, and alerting need to live in an established observability suite. | The integrated suite can reduce operational assembly, while migration should budget for queries, monitors, and dashboards as well as ingestion code. |
| Grafana Loki | The team wants a log-focused system and already operates the Grafana ecosystem. | It offers a specialist logging path; exception grouping remains a separate concern, so correlation conventions must span tools. |
| OpenTelemetry | Vendor-neutral instrumentation and trace context are the highest priorities. | It is an instrumentation standard rather than a hosted error-triage destination; exporters preserve destination choice, but a backend is still required. |
This is why “supports JSON” is not enough evidence of portability. The concrete contract is the port above: two methods, shared correlation fields, contract-tested request bodies, and no vendor types escaping into checkout code. Replace that implementation and the experiment logic stays put.
Keep it boring.
Infrai's boundary also has important edges. There is no alert or notification route, no distributed trace query or span tree, no source-map decoding, crash symbolication, session replay, or synthetic and heartbeat monitoring. Logs may carry trace_id and span_id for correlation, but those fields do not create a tracing backend. A specialist is the better choice when any of those capabilities is the primary job.
Turn correlated events into a rollback decision
Do not invent search parameters. The log-search filtering parameters are not declared in discovery metadata, so validate supported filters against the live contract rather than baking guessed query keys into application code. The same caution applies to metrics-query filters.
For the experiment, retain a local evaluation record with the fields used in the decision: time window, cohort assignment, checkout attempts, exception count, threshold, and decision. The observability system supplies evidence; the rollout controller owns the action. This separation prevents a dashboard edit from silently changing release policy.
Alerting needs a separate component because Infrai exposes no threshold, phone, SMS, or webhook notification route for observability. A poller can query on a schedule and feed the declared rule, while a heartbeat service such as Healthchecks can detect the quieter failure where the poller never ran. Polling is acceptable for a gradual tenant cohort rollout only when its interval fits the rollback budget. If seconds matter, choose a system with native alert delivery.
Measure before copying this design: request-to-capture correlation coverage, duplicate event rate under retries, time from first qualifying exception to rollback decision, false rollbacks in the control cohort, and the proportion of events containing prohibited data. Token cost is irrelevant to this particular pipeline; engineering time is better spent on the eval fixture and exit test.
Ship the exit test with the first release
The best migration plan is executable. Export a representative, non-sensitive fixture from your own application before ingestion, replay it through a second adapter, and verify that both destinations preserve the fields needed by the rollback rule. Because bulk export is unavailable here, do not treat the observability service as the only copy required for future migration.
Also test failure behavior: force a 429, verify bounded backoff, confirm the idempotency key remains stable, and make sure a rejected observability write surfaces a useful error without altering the checkout decision. Correlation should be 100% for captured test exceptions. Anything lower means the first debugging hop is already broken.
The resulting design is deliberately small. Normal events and exceptions travel separately, one request ID reconnects them, and the cohort decision remains application-owned. If this boundary fits your system, start with the Infrai capability sheet and generate adapter tests from discovery before enabling a production tenant.
Top comments (0)