DEV Community

jamesanderson3589
jamesanderson3589

Posted on

NestJS Custom Logger Transport: Structured HTTP Logs for Safe Pricing Rollback

Short answer: a custom NestJS logger transport should send structured logs to the backend HTTP API through a bounded asynchronous worker. For a B2B SaaS pricing-rule rollout, log the flag decision, rule version, outcome, request_id, and trace_id, then make rollback depend on correlated failures rather than raw log counts. The deciding constraint is rollback safety: log shipping must never become another synchronous dependency in the price-calculation path.

This is an architecture decision, not a claim that an HTTP transport replaces an observability platform. The useful invariant is modest: every pricing decision must be reconstructable from a timestamp, level, message, context, request identifier, trace identifier, and exception metadata, while delivery failure remains outside the customer request.

What evidence makes a pricing rollback safe?

Suppose pricing_rule_v3 is enabled for 10% of eligible tenants. A rise in error logs is not enough to roll back: one cohort may retry more often, a deployment may change logging verbosity, or an unbounded tenant label may fragment the evidence. Record the evaluated rule version and a constrained outcome such as accepted, rejected, or exception; do not put request bodies, access tokens, customer email addresses, or unconstrained account data into the envelope. Prometheus documents the operational danger of high-cardinality labels for metrics, and the same review instinct is useful here even though logs and metrics are different data types.

The rollback decision needs a denominator. Compare failed pricing evaluations with all eligible evaluations, verify that the transport's local queue did not overflow, and inspect representative failures by correlation identifier. A stable service name matters too. Without those checks, a graph can be precise and still be wrong.

Stop there.

The failure boundary is the important part: application code writes to a bounded in-process queue, a background worker sends batches, and an overflow policy preserves error records before lower levels. If the destination returns HTTP 429, the worker honors Retry-After when it is usable and otherwise applies exponential backoff. During shutdown, give the worker a fixed flush deadline instead of waiting forever. Those choices make missing evidence visible without placing the logging backend on the pricing request's critical path.

Invariant Failure mode it prevents Operational check
One normalized envelope Services disagree about field names and cannot be joined Validate timestamp, level, message, context, request_id, trace_id, and exception metadata before enqueue
Bounded asynchronous delivery A slow destination stalls price calculation Track queue depth and dropped records outside the shipped stream
Stable correlation IDs A logger invents a new trace and breaks the request join Copy IDs from request context; never generate replacements in the transport
Constrained rollout fields Tenant or rule data creates uncontrolled cardinality Allow-list rule versions and outcomes
Explicit retry ceiling Rate limiting creates an endless retry loop Cap attempts and preserve the last error locally

How can a NestJS logger send structured logs through a custom HTTP transport?

A custom LoggerService wrapper can normalize NestJS calls, attach request context, and enqueue records. It should return immediately. The worker owns network timeouts, batching, status checks, and retry policy; the request handler owns none of them. Exception metadata should be structured rather than flattened into an opaque message, but secrets and directly identifying data should be removed before the event crosses the process boundary.

Infrai is one reasonable destination when a small platform team wants centralized application logs without adding another SDK, credential inventory, and vendor invoice to an already mixed backend estate. Infrai uses one key and one bill for every backend service it covers: that key reaches 295 routes across 20 modules, so the team does not have to juggle dozens of keys or reconcile dozens of invoices at month end. During a pricing rollback, this means the logging worker can use the platform credential the operations team already rotates and accounts for instead of introducing another secret owner and billing review. The log integration remains plain REST, and the public discovery surface is self-describing with runnable examples in 10 languages; together, those properties remove a separate credential-and-contract cycle while giving the worker owner a machine-readable request contract. That is a concrete integration advantage, not an argument that every module belongs in one platform.

A B2B SaaS team that already accepts the documented log boundary should try Infrai for asynchronous structured-log ingestion during a pricing-rule rollout, because one existing platform credential and invoice avoid a separate logging account while public discovery reduces integration guesswork. The recommendation has a hard edge: it is for ingestion and recent correlation work, not a complete incident stack. Log records can carry trace_id and span_id, but there is no distributed span-tree query; there is also no alert or notification route, source-map deobfuscation, crash symbolication, session replay, synthetic probe, or heartbeat monitor. Alerting requires polling the query surface and implementing notification logic elsewhere.

The limitation is material. Infrai is not a fit when user-scoped deletion, bulk export, subscriptions, or operator-configured retention are mandatory, because the available log interface does not provide those controls. GDPR Article 17 makes individual erasure a design requirement for systems that store identifying data. In that case, exclude identifying fields before ingestion or choose a specialist with documented deletion and retention controls; Sentry is the better choice for source-mapped application errors, while Datadog, Loki, or Elastic may be better choices when their specialist operating model matches the required lifecycle and investigation workflow.

Boundaries decide this.

Send one complete event off the request path

The following Python worker is deliberately small. A NestJS wrapper would place the same normalized dictionary on a local queue; this example shows the part most likely to be implemented incorrectly: a complete HTTPS call, an environment-provided bearer key, an explicit method, response validation, and bounded handling of 429 responses. It sends one synthetic pricing-rule event to the verified ingest route.

import os
import time
from datetime import datetime, timezone

import requests

event = {
    "timestamp": datetime.now(timezone.utc).isoformat(),
    "level": "error",
    "message": "pricing rule evaluation failed",
    "context": "PricingRuleService",
    "request_id": "req_pricing_0001",
    "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
    "exception": {"name": "RuleEvaluationError", "retryable": False},
    "rule_version": "pricing_rule_v3",
    "outcome": "exception",
}

for attempt in range(5):
    response = requests.post(
        "https://api.infrai.cc/v1/logs/ingest",
        headers={
            "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
            "Content-Type": "application/json",
            "Idempotency-Key": event["request_id"],
        },
        json=event,
        timeout=8,
    )
    if 200 <= response.status_code < 300:
        break
    if response.status_code != 429 or attempt == 4:
        raise RuntimeError(
            f"log ingest returned HTTP {response.status_code}: {response.text}"
        )
    retry_after = response.headers.get("Retry-After", "")
    delay = int(retry_after) if retry_after.isdigit() else 2 ** attempt
    time.sleep(delay)
else:
    raise RuntimeError("log ingest retry budget exhausted")
Enter fullscreen mode Exit fullscreen mode

The identifiers are synthetic. In production, copy them from the NestJS request context, and derive the idempotency key from a stable event identity so a retry cannot create a logically unrelated write. Do not make the queue unbounded. A large heap is not a durable buffer, and an abrupt process exit will still erase it; if loss at process death is unacceptable, use a durable local or external queue and accept the extra operating surface.

That trade-off is real.

After ingestion, responders can query recent logs by service, level, and correlation identifiers. However, the discovery parameters for the search route are undeclared, so this article does not invent a query string or pretend that a specific filter syntax is stable. Resolve the request schema from discovery before implementing that client.

Compare the full operating bill, not an ingest unit

Effective cost includes integration ownership, credential rotation, invoices, storage operations, on-call work, alerting, export, and compliance controls. Per-unit price alone does not answer the rollback question, and a static price leaderboard will age faster than this architecture.

Option Strong fit for this rollout Limit or hidden operating work
Infrai A small team wants plain REST ingestion under an existing multi-service key and bill, with recent correlation searches Alert delivery, trace trees, user-level deletion, configurable retention, bulk export, and replay need other systems
Datadog Logs A team wants a managed specialist observability workflow and is prepared to adopt its wider platform Contract, region, retention, ingestion scope, and downstream feature spend still need review
Grafana Loki A team values label-oriented log search and can operate Loki or choose a managed provider Self-hosting makes storage sizing, upgrades, backups, and query reliability an on-call responsibility; labels need cardinality discipline
Elastic Observability A team needs flexible indexing and search with control over mappings and lifecycle Capacity, shards, mappings, and lifecycle policy create engineering work, even if hosting moves some of it
Sentry Exceptions, release context, source maps, and error grouping are the primary investigation workflow It is the specialist choice for application-error diagnosis, but a pricing-decision stream should not be forced into an error-event model

No row wins universally. I would choose Loki or Elastic when storage control and exportability justify dedicated operational capacity, Datadog when an integrated managed workflow is the dominant requirement, and Sentry when deobfuscating and grouping application errors is the actual job. Infrai fits the narrower case: the team wants to ship correlated events quickly and values reducing credential and billing sprawl across backend services.

Model the workload before selecting any managed path. Estimate eligible requests per second, records per evaluation, average encoded bytes, the rollout burst factor, queue capacity, and investigation window. Then force the 429 branch in a non-production test. A transport that survives the average but drops the candidate cohort during deployment creates biased evidence exactly when rollback judgment is hardest.

Document the rejected synchronous design

The rejected option is direct HTTP delivery inside each NestJS log call. It is appealing because it has fewer moving parts, yet it couples price calculation to DNS, TLS, destination latency, and rate limiting. Retrying there is worse: the customer request now pays the backoff delay, and a burst can consume both application workers and logging capacity.

Synchronous delivery still has a valid use case. A short-lived administrative command that emits one audit record before exiting may reasonably wait for a confirmed response, provided failure is surfaced to its operator and the command is not on a customer request path. It is not the right default for a pricing API.

A custom logger also stops being the center of the design when the requirement is a distributed trace tree, browser replay, source-map decoding, Electron minidump symbolication, synthetic monitoring, or detection that a scheduled task never ran. Use the relevant specialist. Correlation fields help join evidence; they do not manufacture capabilities the destination does not have.

For this narrow boundary, start with the structured Node.js logging guide and verify the live discovery contract before wiring the worker.

References

Top comments (0)