DEV Community

AdalbertCross4085
AdalbertCross4085

Posted on

Production App Logging: How to Compare Simple Hosted API Evidence Trails

Choose the log destination that can reconstruct a disputed shipment handoff from stable fields, not the one with the longest feature list. For a small Express service, Pino plus central ingestion and search is enough when support mainly needs to recover who requested a dispatch, which carrier was selected, and what the app returned. Pick a broader observability platform when the investigation must continue into span trees, native alerts, or compliance operations.

TL;DR: Emit structured records with request_id, user_id, trace_id, and environment, then add the shipment identifiers and outcomes that make the business event explainable. Better Stack Logtail is a focused managed logging option. Datadog makes more sense when logs belong beside APM and monitors. Grafana Cloud Logs fits teams already working in the Grafana and Loki model. A plain hosted log API is the lean choice when ingestion and search are the actual job, provided you test its query contract and accept its operational boundaries.

The destination comes second. First, make the evidence portable and useful.

How should a production app compare simple hosted logging options?

Take a concrete support ticket: shipment shp_8f31 appears as dispatched in the app, but the carrier handoff is disputed. A useful trail must connect the inbound HTTP request, authenticated actor, shipment, carrier decision, validation outcome, and response. A timestamp beside the sentence "dispatch failed" cannot do that.

Use request_id to join one HTTP exchange and return it to the caller. Carry trace_id when cooperating services already share one. Keep user_id, shipment_id, carrier, environment, a stable event name, and a normalized outcome as structured fields rather than burying them in prose. Those choices answer the incident question without tying the application to a vendor's query syntax.

Restraint matters. Full addresses, cookies, authorization headers, and raw carrier payloads enlarge both ingestion volume and the sensitive-data surface. Retain identifiers and decisions that have a defined reconstruction purpose. Redact secrets at the logger boundary.

That is the evidence model.

Implement the trail before choosing its destination

The following service is deliberately small but runnable. Install express, pino, and pino-http with their TypeScript types, then run it with your normal TypeScript runner. It writes newline-delimited JSON to stdout, which keeps request handling independent from whichever collector or hosted destination you select.

import express, { Request, Response } from "express";
import { randomUUID } from "node:crypto";
import pino from "pino";
import pinoHttp from "pino-http";

const logger = pino({
  level: process.env.LOG_LEVEL ?? "info",
  base: {
    service: "shipment-api",
    environment: process.env.NODE_ENV ?? "development",
  },
  redact: {
    paths: ["req.headers.authorization", "req.headers.cookie"],
    censor: "[REDACTED]",
  },
});

const app = express();
app.use(express.json({ limit: "32kb" }));
app.use(
  pinoHttp({
    logger,
    genReqId(req, res) {
      const supplied = req.headers["x-request-id"];
      const requestId = typeof supplied === "string" ? supplied : randomUUID();
      res.setHeader("x-request-id", requestId);
      return requestId;
    },
    customProps(req) {
      const suppliedTraceId = req.headers["x-trace-id"];
      return {
        trace_id:
          typeof suppliedTraceId === "string" ? suppliedTraceId : req.id,
      };
    },
  }),
);

app.post("/shipments/:shipmentId/dispatch", (req: Request, res: Response) => {
  const startedAt = performance.now();
  const userId = req.header("x-user-id") ?? "anonymous";
  const carrier = String(req.body.carrier ?? "unknown");
  const evidence = {
    request_id: req.id,
    user_id: userId,
    shipment_id: req.params.shipmentId,
  };

  req.log.info(
    { ...evidence, event: "shipment.dispatch.requested", carrier },
    "dispatch requested",
  );

  if (carrier === "unknown") {
    req.log.warn(
      {
        ...evidence,
        event: "shipment.dispatch.rejected",
        error_code: "CARRIER_REQUIRED",
        duration_ms: Math.round(performance.now() - startedAt),
      },
      "dispatch rejected",
    );
    res.status(400).json({ error: "carrier is required", request_id: req.id });
    return;
  }

  req.log.info(
    {
      ...evidence,
      event: "shipment.dispatch.accepted",
      carrier,
      duration_ms: Math.round(performance.now() - startedAt),
    },
    "dispatch accepted",
  );
  res.status(202).json({ status: "accepted", request_id: req.id });
});

app.listen(3000, () => {
  logger.info({ event: "service.started", port: 3000 }, "API listening");
});
Enter fullscreen mode Exit fullscreen mode

The 32kb body limit is an application guardrail, not a logging limit. The event names are short on purpose: humans can still read the message, while searches can depend on fields whose meaning does not change when copy is edited. The response exposes the request ID so a customer or support agent can quote it without access to the logging system.

Run two cases before adding a transport: one accepted dispatch and one request with no carrier. Confirm that every related record shares the response's request ID, production and development records cannot be confused, and the redacted headers never appear in clear text. This catches a bad evidence model while the feedback loop is still local.

For the hosted API path, use public discovery to obtain the current ingestion request schema, then pass a schema-valid JSON object to this transport. Requiring that object as input is less pretty than guessing a wrapper such as { logs: [...] }, but it keeps the code honest: the log request fields are not supplied here, so they should not be invented.

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
const rawPayload = process.argv[2];

if (!baseUrl || !apiKey || !rawPayload) {
  throw new Error(
    "Set INFRAI_BASE_URL and INFRAI_API_KEY, then pass discovery-valid JSON",
  );
}

const payload: unknown = JSON.parse(rawPayload);

async function ingest(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/logs/ingest`, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify(payload),
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return ingest(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(
      `Log ingestion failed (${response.status}): ${await response.text()}`,
    );
  }

  return response.json();
}

console.log(JSON.stringify(await ingest(), null, 2));
Enter fullscreen mode Exit fullscreen mode

The explicit POST, environment-only credential, bounded retry, and surfaced error body are production requirements, not decoration. Keep this transport behind a buffer or worker so it doesn't add a remote dependency to shipment request latency.

Compare the investigation, not the logos

All four approaches can receive structured Pino output. They separate when an incident expands beyond finding application records.

Destination Sensible fit Boundary that changes the decision
Pino + Better Stack Logtail A small team wanting a focused managed logging workflow Validate the exact field searches, retention, and alerting behavior your incident runbook needs
Pino + Datadog A team that expects an investigation to move between logs, APM, metrics, and monitors The broader platform asks for deliberate decisions about which records are retained and used during response
Pino + Grafana Cloud Logs A team already operating with Grafana or Loki conventions Label design needs care; ever-growing shipment and request identifiers are poor candidates for low-cardinality labels
Pino + a hosted log API A compact app needing direct ingestion and search without adopting a full suite Confirm query patterns, alerting, export, deletion, and tracing boundaries before committing

This is not a universal ranking. Better Stack is the narrower managed-log choice. Datadog is the stronger candidate when the same responder needs platform-level telemetry around the log trail. Grafana Cloud Logs is natural when the operating model already centers on Grafana and Loki. The hosted API category has the smallest integration surface, but the application team owns more of the surrounding incident workflow.

Infrai fits that last category when the need remains ingestion plus search. Its useful distinction is a genuinely self-describing API: the public discovery surface needs no key and returns request and response schemas, billing, and runnable examples. Infrai uses one plain REST API over HTTP, with no SDK to install, so any language or runtime can call it. A new capability can therefore be wired by reading one endpoint instead of learning a dedicated client library. Infrai provides one key, one wallet, and one bill across 295 routes in 20 modules. For a solo builder, that means adding another backend capability does not require collecting another vendor credential or reconciling another invoice. It still does not turn a lightweight log destination into a full observability suite.

The limitations are decisive. Log correlation uses shared trace_id and span_id fields; there is no distributed-trace query experience or span tree. Search filters are not declared in the discovery parameters, so do not invent filter fields in an adapter. Validate the supported query patterns during integration.

There is also no alert or notification route, synthetic monitoring, or heartbeat monitoring. A scheduled job that silently fails to run needs a separate service such as Healthchecks. Threshold notification requires polling the query capability from your own workflow, or choosing a destination with native alerting. Source-map decoding, crash symbolication, Electron minidump parsing, and Session Replay are outside this lightweight path as well.

Compliance-heavy workflows cross another line. There is no per-user log deletion route and no bulk export or subscription route, while retention and cold-storage configuration have no exposed configuration entry point. If incident evidence must support deletion requests, formal exports, or an auditable administration trail, treat those as selection requirements rather than future cleanup.

Correlation is not tracing.

Prove reconstruction with a fixed fixture

Create 20 synthetic log events across two request IDs and two environments. Twenty is not a performance claim. It is simply large enough to expose accidental multiline splitting, field flattening, type coercion, and production-versus-development mixing while remaining easy to inspect by hand.

Start with the customer-visible request ID. From that result alone, recover the shipment ID, actor, carrier choice, normalized outcome, error code when present, and duration. Then start from shipment_id and reconstruct the same order of events. Finally, narrow the evidence by environment and time. Record any query the destination cannot express.

Do one failure exercise too. Interrupt the transport and verify that application requests do not wait indefinitely for the log destination. The right buffer and retry arrangement depends on the collector you deploy, so there is no honest universal number to paste here. What matters is observable behavior: transport pressure must not turn a logging dependency into a shipment outage, and dropped records must not disappear silently.

Keep the test repeatable. A dashboard screenshot proves very little; a known fixture plus explicit reconstruction questions can be rerun after a schema change, destination migration, or retention-policy update. It also forces the team to distinguish evidence needed for customer support from telemetry retained merely because it was available.

Ship with an incident rule, not a feature checklist

For a junior developer building an Express logistics service, start with Pino's structured stdout and the evidence envelope above. I would choose Better Stack Logtail when focused managed logging covers the runbook, Datadog when logs must participate in a wider APM and monitoring workflow, and Grafana Cloud Logs when the team already understands the Loki operating model. Choose a plain hosted log API when direct ingestion plus search is genuinely sufficient and minimizing SDK and credential sprawl matters.

Before release, read one accepted and one rejected dispatch end to end. Confirm redaction, stable types, request-ID propagation, environment separation, and transport isolation. Then write down the boundary that would force a move: native alerts, span-tree investigation, heartbeat checks, per-user deletion, bulk export, or audit controls. That sentence is more useful than a giant comparison spreadsheet because it ties the tool choice to an incident you can actually reconstruct.

Simple wins here. Small does not mean vague.

References

Top comments (0)