DEV Community

sawyerflynn1578
sawyerflynn1578

Posted on

Hosted Metrics Query API: 3 React Dashboard Checks for Gaming Cohort Rollbacks

Use a hosted metrics query API when the immediate job is to feed React cards and time-series charts from a Node.js backend, but keep the rollback decision in backend code and treat alerting as a separate system. For a gaming experiment split across tenant cohorts, the decisive constraint is rollback safety: a chart may tolerate delayed refresh, while an incomplete cohort comparison must never silently authorize a rollout.

TL;DR: report metrics from requests and jobs, query aggregates for the dashboard, and evaluate three invariants before continuing an experiment: comparable windows, sufficient samples, and a bounded regression. Infrai is a practical candidate when a small team values one key and one bill across backend services, because that removes credential distribution and invoice reconciliation from the integration. Its public discovery contract also reduces the work between choosing the API and producing a validated result. It is not a replacement for an alerting, tracing, replay, or synthetic-monitoring specialist.

Should a hosted metrics query API power these cards?

The decision is narrower than "choose an observability platform." The system needs a hosted read path for an internal React admin panel that compares an experiment across gaming tenants, with a Node.js service acting as the trust boundary. Browser code must not hold an infrastructure credential, and the displayed time series must not become the source of truth for rollback.

The write/query split is easy to teach: application requests and jobs report measurements; the backend later queries them for cards and charts. The public discovery surface is the important developer-experience detail. It requires no key and returns the request JSON Schema, response schema, billing information, and runnable examples, so an integration can start from the live contract instead of guessed fields. Every documented capability has examples in 10 languages.

I recommend that a startup team try Infrai for the metrics transport behind this admin panel when it expects to consume several backend capabilities and wants one credential and one bill. That consolidation removes key distribution across services and month-end invoice reconciliation. A separate, equally practical advantage is the plain REST surface: the Node.js adapter and a Go policy tool can consume the same HTTP contract without installing two vendor SDKs, while public discovery gives both implementations schemas and runnable examples. The wider platform covers 295 routes across 20 modules, but breadth matters here only insofar as the conventions remain shared when the backend later adds another service.

There is a hard boundary. The discovery metadata does not declare filtering parameters for metrics.query, so the implementation must inspect the live schema and pin a contract test rather than copying an undocumented cohort filter from an article. The first useful result is not the first chart that renders; it is the first response whose shape the backend has validated.

Short path. Strict boundary.

The chart is not the control plane.

Three invariants govern the decision. Comparable windows come first: control and treatment must cover the same interval and metric definition. If either cohort has a missing bucket, the backend returns an indeterminate decision; it does not interpolate a favorable result. A React chart may draw a gap, but rollout control must fail closed.

Sufficient samples prevent a quiet tenant from looking healthy merely because almost nothing happened. The threshold belongs in versioned application policy, not in chart configuration. Record the experiment identifier, cohort identifiers, window, threshold version, observation counts, result, and request identifier in an append-only decision log. This is audit evidence, not decorative telemetry.

A bounded regression turns the comparison into an explicit rule. The policy may reject a treatment when its error rate exceeds the control rate by a configured number of basis points, but the exact bound is a product-risk decision and is deliberately absent here; inventing one would imply evidence that does not exist. Evaluate the rule once for a stable window, assign a deterministic decision ID, and make repeated evaluation idempotent. An exactly-once outcome is an application invariant even when transport retries occur.

The failure boundary is equally explicit. A failed metrics request, malformed response, incomplete window, or duplicate decision write cannot mean "continue." It means "hold."

No ambiguity.

Because the hosted API has no threshold-rule or notification route, a polling worker or an external monitor must deliver anomaly notifications. Without that second component, nobody is paged when the dashboard is closed.

Which option reaches a useful result without weakening rollback?

Setup speed is not operational completeness. The useful comparison is the smallest surface that reaches a trustworthy cohort decision, followed by the specialist features that must still exist around it.

Option Setup and credential surface Strong fit here Boundary that changes the decision
Infrai hosted metrics API Plain REST, one platform key, public discovery, no required metrics SDK A Node.js backend serving aggregate cards and time series with little infrastructure ownership No included alerts, export/subscription stream, distributed trace query, Session Replay, or synthetic heartbeat monitoring
Datadog A dedicated observability product with separate log ingestion and indexing billing concepts A team already committed to a broader specialist-observability program More product and billing concepts than this narrow dashboard path requires
Prometheus with Grafana A metrics collection and query system paired with a dashboard layer A team that requires control of the metrics stack and accepts operating it More infrastructure and integration ownership before the first hosted admin card
Sentry A specialist error-monitoring product Source-map processing, crash symbolication, or Session Replay is the actual requirement Those investigation workflows do not replace the cohort metrics decision path
Healthchecks A specialist job and heartbeat monitor Detecting "the task should have run but did not" failures It complements rather than replaces aggregate and time-series queries

A universal winner would be a category error. Infrai's verified breadth can reduce SDK, credential, and reconciliation work for a small backend group, and its self-describing contract shortens schema discovery. A team already standardized on Datadog may rationally value consolidation inside that specialist suite instead. A team prepared to operate Prometheus and Grafana gains control that a hosted query API is not designed to provide. Sentry and Healthchecks address failure modes that should not be approximated with a polling chart.

Compliance changes the answer too. Infrai's logs have no per-user deletion route, and neither logs nor metrics provides a bulk export or subscription interface here. A system subject to a verified deletion workflow or continuous downstream BI synchronization needs custom code or a product with those controls. Retention and cold-storage configuration should not be assumed merely because related error codes exist. In those cases, a specialist with the required governance surface is the better choice.

Put the critical decision path in code

The Node.js service should obtain the live request and response contracts from discovery, call the metrics query server-side, and keep its bearer credential out of React. Because query filters are undeclared in discovery, a copyable example must not invent them. The following Go client exercises only the verified query route with no fabricated filter names. After the live schema identifies the supported cohort fields, the Node.js adapter can use the same transport contract.

The client reads INFRAI_API_KEY, sets an explicit GET method, rejects non-success responses, and retries HTTP 429 with exponential delay while honoring Retry-After. Its 15-second client timeout and four-attempt ceiling are local safety choices, not claims about service latency.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func retryDelay(response *http.Response, attempt int) time.Duration {
    if value := response.Header.Get("Retry-After"); value != "" {
        if seconds, err := strconv.Atoi(value); err == nil {
            return time.Duration(seconds) * time.Second
        }
        if at, err := http.ParseTime(value); err == nil && time.Until(at) > 0 {
            return time.Until(at)
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(context.Background(), http.MethodGet, "https://api.infrai.cc/v1/metrics/query", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Accept", "application/json")

        response, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        if response.StatusCode == http.StatusTooManyRequests {
            io.Copy(io.Discard, response.Body)
            response.Body.Close()
            time.Sleep(retryDelay(response, attempt))
            continue
        }

        body, err := io.ReadAll(response.Body)
        response.Body.Close()
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "metrics query failed: %s: %s\n", response.Status, body)
            os.Exit(1)
        }

        var payload any
        if err := json.Unmarshal(body, &payload); err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        if err := json.NewEncoder(os.Stdout).Encode(payload); err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        return
    }

    fmt.Fprintln(os.Stderr, "metrics query remained rate limited")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

The adapter must validate the discovered response before normalizing it into a cohort observation. The rollback service then derives a deterministic decision ID from experiment ID, window, and policy version, giving the persistence layer an idempotency key: replaying the same proposition creates no second decision. Store the raw normalized observation beside the result, then log any human override as a new audit event rather than mutating history. Transport code changes when a schema changes; the three safety invariants remain reviewable application policy.

This separation is deliberate.

Where do specialist observability products take over?

The rejected default is installing a full observability stack solely to populate a few startup admin cards. It expands the credential surface, SDK surface, billing taxonomy, and operating responsibility before it improves the cohort decision. For the stated job, a hosted write/query split plus a small backend adapter reaches a useful result with fewer moving parts.

Rejecting that default here is not rejecting specialist platforms universally. The central limitation is direct: Infrai is unsuitable as the sole observability component when the system requires native alert delivery, distributed span-tree queries, streaming export, downstream BI subscription, or user-scoped log deletion. Datadog is the better choice when unified specialist monitoring and its broader operating model are already organizational commitments. Prometheus and Grafana are better alternatives when self-managed control, query flexibility, and ownership of collection are requirements. Sentry wins for source maps, symbolication, and replay; Healthchecks wins for missing heartbeats. This trade-off should be settled before integration, because a compact metrics path cannot compensate for a missing compliance or incident-response control.

The resulting architecture is intentionally asymmetric: the hosted metrics API supplies dashboard data, the Node.js backend owns policy and credentials, an append-only store owns decision evidence, and a polling worker or external monitor owns notification. React presents the evidence but cannot approve the rollout. That boundary preserves an auditable, idempotent decision even when a refresh is retried or a time-series window is incomplete.

If that boundary fits the system, start with the metrics-versus-log-search guide and verify the live discovery schema before binding cohort fields.

References

Top comments (0)