DEV Community

Haelion14
Haelion14

Posted on

Marketplace Production Alerts: Combine Failure Metrics, Logs, Request and Trace IDs

TL;DR: Page only when a pricing-rule cohort breaches its user-facing failure SLO in one region, then attach a correlation key that lets the responder move from the alert to logs and traces. Keep raw errors as investigation evidence, not separate pages. A polling worker can evaluate the rule every 30 seconds, but it needs bounded labels, delayed-data tolerance, and an explicit rollback gate.

For a marketplace releasing a new pricing rule behind a flag, the useful question is not whether the application emitted an error. It is whether flagged quote requests in the EU or US are failing often enough, for long enough, to threaten the rollout objective. That distinction removes noise while preserving the evidence needed to explain a bad price, a timeout, or a downstream rejection.

How should production failure alerts combine metrics, logs, and trace IDs?

Start with the customer outcome. Define a pricing evaluation as successful only when the request returns a valid quote under the intended rule version; count timeouts, invalid results, and dependency failures against that outcome. Alert on a sustained burn of the error budget, split by region and flag cohort, rather than on each exception class. OpenTelemetry semantic conventions provide common telemetry attribute names, while W3C Trace Context defines interoperable traceparent propagation. Those standards make correlation portable. They do not decide what deserves a page.

The signal should answer four things: which SLO is burning, which region is affected, whether the flagged cohort differs from control, and which rule version is involved. Request IDs remain useful for support searches and application logs, but trace IDs are the cross-service correlation key. Do not turn either identifier into a metric label. Prometheus warns that every unique label combination creates a new time series and cautions against unbounded-cardinality labels.

Low-volume regional cohorts are a trap. Five failures can look dramatic as a percentage yet provide weak evidence; meanwhile, a global aggregate can hide a genuine EU-only regression behind healthy US traffic. Require both a minimum event count and an error-budget burn condition.

Quiet is a feature.

Build the correlation path before the pager

At ingress, accept a valid W3C trace context or create a new trace, then create a separate request ID if the application needs an opaque lookup key. Propagate trace context through the pricing service, queue messages, and downstream calls. Log the trace ID, request ID, region, flag cohort, rule version, outcome, and error class as structured fields. Record metrics with only bounded dimensions such as region, cohort, outcome, and rule version; never include request ID, trace ID, user ID, listing ID, or raw error text.

Each signal gets one job. Metrics decide that a symptom is broad enough to page. Traces show the critical path and dependency timing. Logs retain detailed decision context that would be wasteful or unsafe as metric labels. An exception event belongs on the active trace when possible, with the structured log carrying the same trace ID for searches outside the sampled trace set.

This evaluator core uses a generic query boundary and aggregated windows. Storage-specific query syntax stays outside the worker.

package alerting

import (
    "context"
    "fmt"
    "time"
)

type Window struct {
    Region, Cohort, RuleVersion string
    Attempts, Failures          uint64
    Exemplars                   []string
}

type QueryClient interface {
    PricingWindows(context.Context, time.Time, time.Time) ([]Window, error)
}

type Notifier interface {
    Page(context.Context, string, string, []string) error
    Resolve(context.Context, string) error
}

type Worker struct {
    Query QueryClient
    Notify Notifier
    MinAttempts uint64
    ErrorRate float64
    Lookback, DataDelay time.Duration
}

func (w Worker) Poll(ctx context.Context, now time.Time) error {
    end := now.Add(-w.DataDelay)
    windows, err := w.Query.PricingWindows(ctx, end.Add(-w.Lookback), end)
    if err != nil {
        return fmt.Errorf("query pricing windows: %w", err)
    }
    for _, v := range windows {
        key := fmt.Sprintf("pricing:%s:%s:%s", v.Region, v.Cohort, v.RuleVersion)
        firing := v.Attempts >= w.MinAttempts &&
            float64(v.Failures)/float64(v.Attempts) >= w.ErrorRate
        if firing {
            summary := fmt.Sprintf("pricing SLO burn: region=%s cohort=%s rule=%s failures=%d attempts=%d",
                v.Region, v.Cohort, v.RuleVersion, v.Failures, v.Attempts)
            if err := w.Notify.Page(ctx, key, summary, v.Exemplars); err != nil {
                return fmt.Errorf("page %s: %w", key, err)
            }
            continue
        }
        if err := w.Notify.Resolve(ctx, key); err != nil {
            return fmt.Errorf("resolve %s: %w", key, err)
        }
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The production wrapper should apply a timeout shorter than the polling interval, retry transient reads with bounded backoff, and persist alert state so restarts do not send a fresh notification for every still-firing cohort. Use a stable deduplication key. Cap exemplar trace IDs to a fixed count, redact sensitive marketplace fields before export, and retain only telemetry required by incident and compliance policies.

Do not interpret a failed telemetry query as proof that pricing is healthy. Emit a separate health signal for the evaluator, and route repeated evaluation failures as an observability-path problem. Mixing that condition into the pricing alert makes ownership ambiguous during an incident.

Set thresholds from an SLO, not intuition

A threshold such as ErrorRate: 0.02 in an example is not universal guidance. Derive the trigger from the pricing SLO, its evaluation window, normal request volume, and the error budget the team will consume before intervention. Multi-window burn-rate alerting helps because a short window detects fast damage while a longer window rejects brief spikes; Google's SRE Workbook documents this pattern and its trade-offs.

Capacity planning belongs here. Estimate metric series before rollout as the product of bounded dimensions: two regions, two cohorts, three outcomes, and two simultaneously active rule versions produce 24 series for one counter family. Adding a request ID turns that finite budget into traffic-shaped growth. Estimate trace and log ingestion separately from peak requests per second, sampling policy, average event size, retention, and regional residency requirements. Averages conceal launch-hour load.

The buy-versus-build choice is operational, not ideological.

Concern Build the evaluator Use a managed evaluator Decision evidence
Rule control Exact cohort and SLO semantics Supported alert model sets constraints Can it express minimum volume plus burn rate?
On-call load Team owns state, retries, upgrades, and capacity Provider owns more control-plane work Which failure domains still page the team?
Portability Generic interfaces require engineering Query boundaries can create lock-in Can telemetry and alert state migrate?
Regional handling Placement and retention are team work Available regions set boundaries Do EU and US residency needs match?
Cost shape Infrastructure, engineering, and on-call time Usage, integration, and exit work What happens at peak volume and longer retention?

No row decides the answer alone. A managed service can reduce control-plane work while leaving instrumentation, SLO design, and incident response with the team; self-hosting can improve control while adding upgrade and capacity obligations. Put both into the roadmap calculation.

Verify the rollout and rehearse rollback

Before exposing the flag, send synthetic successful and failed evaluations through each region. Verify that metrics contain bounded labels, logs carry both correlation identifiers, and a trace crosses every asynchronous boundary. Then inject enough failures into a test cohort to cross the configured minimum volume and burn threshold. Confirm one deduplicated page arrives with representative trace IDs, not one page per exception.

Next, test absence. Delay telemetry beyond the worker's allowance, make the query dependency unavailable, restart the worker while an alert is firing, and remove traffic from one cohort. Each case needs a known state: delayed evaluation, evaluator-health alert, preserved deduplication, or insufficient-volume suppression. Unknown must never silently become healthy.

Rollback needs its own gate. Pause flag expansion when the fast burn alert fires; disable the pricing rule for the affected region and cohort through the existing flag control; then verify that new requests use the prior rule version and both short and long SLO windows recover. Keep the alert open until the recovery window is satisfied. A falling exception count is inadequate because traffic may also have fallen.

I would require this evidence in the change record: flag allocation by region, the SLO and minimum-volume values, links to the metric view and correlated traces, the rollback owner, and the exact condition for resuming exposure. This is an explicit trade: slower expansion in exchange for a decision trail the on-call engineer can trust.

The operating rule

A production alert should represent threatened marketplace behavior, not telemetry activity. Page from bounded, cohort-aware SLO signals; investigate through trace IDs and structured logs; and treat the polling worker as a monitored stateful control component. The result is fewer alerts, but each carries enough regional and rollout context to support a defensible rollback decision.

References

Top comments (0)