DEV Community

MerrickVance8452
MerrickVance8452

Posted on

React Error Boundary Reports — Send Frontend Logistics Failures to Your Backend

Short answer: send exceptions from a React error boundary, window.onerror, and unhandledrejection to a small backend endpoint, but choose this design only when basic runtime evidence is enough. For a nightly logistics pipeline, the useful invariant is not which error vendor stores the event; it is that every report carries the release, page URL, browser, an appropriate user identifier, and a client-generated fingerprint. That stable contract lets the storage provider move without forcing a frontend rewrite. It does not turn a log-shaped event into polished crash analysis.

Consider a bounded incident: the nightly shipment import completes, but the operations screen crashes while rendering one malformed consignment. The batch logs say the run succeeded. The browser knows which release, route, and stack failed, while the backend knows which import produced the data. Without a shared correlation value, reconstructing the sequence becomes guesswork. I would capacity-plan this path as incident evidence, with its own delivery budget and SLO, rather than treat it as an unlimited duplicate of every console message.

One warning deserves the opening page: minified stacks remain minified because this lightweight path has no source-map deobfuscation or symbolication. It also has no session replay. If those are requirements, stop here and use a dedicated crash-analysis product.

How should a React error boundary send frontend JavaScript failures?

Capture three browser failure surfaces: the nearest React error boundary for render failures, the global error event for uncaught runtime errors, and unhandledrejection for rejected promises that escape application handling. Normalize them into one small envelope before posting to your backend. Do not rely on the boundary alone; it does not cover every failure surface.

The envelope should include the page URL, release, browser, user ID only where appropriate, and a client-generated fingerprint. For this logistics workflow I would also carry an existing pipeline correlation value in the application context, then preserve it through storage; Infrai logs can hold trace_id and span_id for correlation, although there is no distributed-trace query or span tree. Do not invent search filters around those fields, because the discovery parameters for log search are undeclared.

The fingerprint is the grouping contract. Derive it from stable inputs such as exception type, normalized top stack frame, and release family rather than the full message, which may contain shipment IDs and create one group per consignment. Keep the raw payload bounded. A 1% frontend failure budget becomes meaningless if one recursive render error can emit thousands of near-identical records and crowd out the event that started the sequence.

Small is safer.

The trade-off is explicit.

This is also where privacy changes the architecture. Logs have no per-user deletion API, so they are a poor fit when GDPR deletion by user is mandatory. Avoid unnecessary PII, minimize the identifier, define retention outside the event contract, and select a system with a suitable deletion workflow if erasure is a hard requirement.

A backend boundary that can survive a provider change

The browser should post to an endpoint you own. That endpoint validates size and content type, removes fields outside the allowlist, applies rate limits, and translates the stable internal envelope to whichever store sits behind it. Browser code never receives an infrastructure API key.

The following Go handler shows that preventative boundary. It is runnable with the standard library, accepts a compact JSON event, computes a fallback fingerprint when the client omitted one, and writes newline-delimited JSON to standard output. In production, replace the encoder with a narrow adapter for the selected provider; keep BrowserError unchanged.

package main

import (
    "bytes"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type BrowserError struct {
    Message     string `json:"message"`
    Stack       string `json:"stack,omitempty"`
    URL         string `json:"url"`
    Release     string `json:"release"`
    Browser     string `json:"browser"`
    UserID      string `json:"user_id,omitempty"`
    Fingerprint string `json:"fingerprint,omitempty"`
    TraceID     string `json:"trace_id,omitempty"`
}

func main() {
    logger := log.New(os.Stderr, "browser-errors ", log.LstdFlags)
    http.HandleFunc("/browser-errors", func(w http.ResponseWriter, r *http.Request) {
        if r.Method != http.MethodPost {
            http.Error(w, "method not allowed", http.StatusMethodNotAllowed)
            return
        }
        if !strings.HasPrefix(r.Header.Get("Content-Type"), "application/json") {
            http.Error(w, "content type must be application/json", http.StatusUnsupportedMediaType)
            return
        }

        r.Body = http.MaxBytesReader(w, r.Body, 32<<10)
        dec := json.NewDecoder(r.Body)
        dec.DisallowUnknownFields()

        var event BrowserError
        if err := dec.Decode(&event); err != nil {
            http.Error(w, "invalid event", http.StatusBadRequest)
            return
        }
        if event.Message == "" || event.URL == "" || event.Release == "" {
            http.Error(w, "message, url, and release are required", http.StatusBadRequest)
            return
        }
        if event.Fingerprint == "" {
            sum := sha256.Sum256([]byte(event.Release + "\x00" + event.Message))
            event.Fingerprint = hex.EncodeToString(sum[:12])
        }

        if err := sendInfrai(r, event); err != nil {
            logger.Printf("capture event: %v", err)
            http.Error(w, "event unavailable", http.StatusServiceUnavailable)
            return
        }
        w.Header().Set("Content-Type", "application/json")
        w.WriteHeader(http.StatusAccepted)
        fmt.Fprintln(w, `{"accepted":true}`)
    })

    logger.Fatal(http.ListenAndServe(":8080", nil))
}

func sendInfrai(r *http.Request, event BrowserError) error {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return fmt.Errorf("INFRAI_API_KEY is required")
    }
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    if baseURL == "" {
        return fmt.Errorf("INFRAI_BASE_URL is required")
    }
    payload := struct {
        ReceivedAt time.Time    `json:"received_at"`
        Error      BrowserError `json:"error"`
    }{time.Now().UTC(), event}
    body, err := json.Marshal(payload)
    if err != nil {
        return err
    }

    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(r.Context(), http.MethodPost,
            baseURL+"/errors/capture", bytes.NewReader(body))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        resp, err := client.Do(req)
        if err != nil {
            return err
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 8<<10))
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            return fmt.Errorf("capture failed: status=%d body=%s", resp.StatusCode, responseBody)
        }
        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-r.Context().Done():
            return r.Context().Err()
        }
    }
    return fmt.Errorf("capture retries exhausted")
}
Enter fullscreen mode Exit fullscreen mode

A real deployment still needs authentication appropriate to the application, origin checks, rate limiting, and a queue or other bounded buffer between intake and storage. Set an explicit payload ceiling and decide what happens under backpressure. My default is to shed repeated fingerprints before unique ones, because incident reconstruction values breadth of evidence more than the 5,001st copy of the same render failure. That is a design rule, not a measured threshold. My first design instinct would be to preserve every report; the capacity-planning correction is that duplicates consume the same finite incident window as novel evidence, so overload policy belongs in the design before launch.

If the adapter sends to Infrai, it can use POST /v1/errors/capture with Authorization: Bearer $INFRAI_API_KEY; keep that credential on the server, check non-success responses, and back off on HTTP 429 while honoring Retry-After. The platform's public discovery describes request schemas, so the adapter can follow the declared contract. The broader benefit is substitution: the browser contract stays fixed while the provider behind it changes.

Buy or build for the failure analysis you actually need

A custom collector and an error-analysis service solve different problems. The former gives control over the intake contract. The latter earns its operational cost by turning opaque crashes into navigable evidence. Evaluate that distinction before debating storage price.

Option Best fit for this logistics workflow Boundary to test before adoption
Custom Go intake plus logs A small event vocabulary, internal correlation, and a team willing to own grouping, retention, and alerting Minified stacks remain difficult; user-level deletion, source maps, replay, and polished triage are absent unless separately built
Sentry The team needs a dedicated error-tracking workflow and should evaluate source-map and replay documentation directly Adds a specialized product contract and its operating model
Bugsnag Release-oriented stability analysis is central and the team prefers a dedicated crash product Validate browser support, grouping behavior, deletion workflow, and export needs against current documentation
Rollbar The team wants managed error triage rather than a general event stream Validate source-map processing, retention, and alert behavior for the exact plan
Datadog Error Tracking Browser errors must sit beside a broader managed observability estate The suite's breadth can be more coupling than a narrow pipeline needs
Grafana The team already operates a Grafana-centered log workflow and wants browser evidence beside it Error grouping and crash-analysis behavior must be assembled and tested for this workflow
Infrai error capture One plain REST contract and one key across backend capabilities matter, while basic capture is sufficient No source-map deobfuscation, symbolication, session replay, or user-level log deletion; alerts require polling and external notification logic

That table is deliberately a shortlist, not a winner. Sentry, Bugsnag, Rollbar, Datadog, and Grafana should be tested with the same minified production fixture and the same deletion request. A polished demo with an unminified stack proves little. Infrai is credible when provider substitution and a consistent REST surface carry more roadmap weight than advanced crash analysis; its live discovery covers 295 routes across 20 modules, but breadth does not erase the error-tracking limitations. It is not a fit when source maps, replay, symbolication, or managed notifications are required; choose a dedicated alternative after verifying those features against its current documentation.

Alerting changes the on-call calculation too. The lightweight path has no threshold, phone, SMS, or webhook notification route. Polling a query API and building notification state is possible, but then the platform team owns deduplication, missed-poll recovery, and alert SLOs. Silent nightly-job failure needs a heartbeat monitor such as Healthchecks because browser exceptions only exist when somebody loads the broken screen.

When should this advice be rejected?

Reject the custom path when responders need readable minified stacks without reproducing locally, replay of the user's session, native crash symbolication, or managed alert delivery. Reject logs as the primary store when per-user erasure is compulsory. Reject browser errors as proof that the nightly pipeline ran: a missing event cannot distinguish success, no traffic, and a dead collector.

Use the lightweight path when the question is narrower: did this release produce a runtime failure on this route, which stable fingerprint groups it, and which pipeline correlation value helps reconstruct the event sequence? Establish an ingestion SLO, send a synthetic canary error after each release, and capacity-plan for a burst caused by a bad deployment. Keep enough headroom that one release can fail noisily without consuming the entire incident window.

I would make the buy decision after a production-like bake-off: minify the same failing bundle, send it through each candidate, request deletion for one test user, disconnect notification delivery, and time the reconstruction steps without pretending that vendor dashboards remove the need for a runbook. The evidence should decide. For basic runtime capture, the small contract is defensible; for crash analysis, buy the analysis.

Sources

Top comments (0)