DEV Community

MerrickVance8452
MerrickVance8452

Posted on

React Dashboard Query API — Hosted Metrics for Tenant Experiment Attribution

TL;DR: Choose a hosted metrics query API only after proving that every experiment series can carry a stable tenant-cohort identifier, a bounded time range, and an attributable ingestion and retention cost. Put a Node.js boundary between React and the provider, return one small response contract for cards and charts, and reject unbounded queries. A convenient query endpoint cannot repair ambiguous cohort membership or uncontrolled label cardinality.

For a B2B SaaS admin panel, the operational decision rule is blunt: if finance cannot map the experiment's telemetry volume to a cohort, and the on-call engineer cannot explain the card's denominator, the integration is not ready. Evaluate the data model first, then the query API. This reverses the usual demo-led selection process, but it protects the two things the demo hides: allocation error and capacity growth.

How should a React dashboard query a hosted metrics API?

A card such as “successful exports” looks like one number. The underlying question is conditional: successful exports for which experiment variant, among which tenant cohort, during what interval, under which definition of success? A time series adds a step and aggregation interval. If cohort membership changes during the experiment, joining current account metadata at read time can silently rewrite the past. Imagine a tenant assigned to the control cohort on Monday and the treatment cohort on Thursday. A Friday query that joins every historical sample to Friday's assignment would move Monday's exports into treatment, changing both cohorts without a single metric sample changing. The chart can be internally consistent and still answer the wrong question. That failure is especially hard to spot when a card shows only the latest total and the detailed series lives on another screen.

The denominator matters.

Record the experiment and cohort dimensions when the measurement is produced, or preserve an immutable assignment record that the query layer can join by event time. The former makes reads simpler but adds labels; the latter limits metric dimensions but requires a trustworthy temporal join. This is the first real trade-off. There is no provider toggle that chooses correctly for the business.

Cardinality is the capacity-planning trap. Prometheus documentation warns that every unique label set creates a new time series and specifically advises against high-cardinality labels such as user IDs. A tenant ID has the same dangerous shape when tenant count grows. Prefer bounded labels such as experiment, variant, cohort, region, and result; keep raw tenant identity in an assignment ledger or exemplars only when the chosen system and privacy policy support that use. Then write the worksheet before opening a trial account: enumerate each bounded dimension, multiply the possible values, add replica and environment dimensions that the collector actually emits, and repeat the calculation for every instrument. Compare that upper bound with the candidate's documented limits and billing dimensions. The worksheet will be imperfect because not every combination occurs, but an explicit overestimate is easier to challenge than a bill explained after the fact.

Suppose planning inputs are 12 experiments, 3 variants, 8 cohorts, 4 regions, and 2 results. Their full cross-product is 2,304 possible series per instrument before replicas or deployment labels. That is arithmetic, not a benchmark. It is useful because adding a tenant_id dimension for 5,000 tenants changes the upper bound by three orders of magnitude.

Stop there.

That is enough to reject the label design before it reaches production.

Define the query boundary before comparing hosts

React should ask the Node.js backend for a product-level answer, not send a provider query language from the browser. The backend owns authentication, authorization, cohort definitions, maximum range, step selection, retries, and the translation from provider responses into a stable contract. This also prevents browser credentials from becoming the access-control model.

Keep it narrow.

A useful contract has fewer degrees of freedom than the upstream API. For example, accept a known experiment, cohort, metric key, start, end, and requested resolution; derive the provider expression on the server. Return timestamps, nullable values, units, and freshness. A missing point is not zero. Zero means the instrument observed a zero; null means the system lacks a value for that bucket.

The validation logic below is Go because the important artifact is the boundary, not a vendor SDK. A Node.js service should enforce the same invariants before it calls its adapter.

package query

import (
    "errors"
    "time"
)

type Request struct {
    Experiment string
    Cohort     string
    Metric     string
    Start      time.Time
    End        time.Time
    Step       time.Duration
}

func (r Request) Validate(now time.Time) error {
    if r.Experiment == "" || r.Cohort == "" || r.Metric == "" {
        return errors.New("experiment, cohort, and metric are required")
    }
    if !r.Start.Before(r.End) || r.End.After(now.Add(5*time.Minute)) {
        return errors.New("invalid time range")
    }
    if r.End.Sub(r.Start) > 30*24*time.Hour {
        return errors.New("range exceeds 30 days")
    }
    if r.Step < time.Minute || r.Step > 24*time.Hour {
        return errors.New("step must be between 1 minute and 24 hours")
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The numbers in that validator are example policy inputs, not universal defaults. Set them from the admin panel's actual decisions: a 15-minute rollout view and a 90-day quarterly review need different routes or pre-aggregations. Letting one endpoint serve both usually produces either wasteful short-range queries or misleading long-range detail.

Cache only after authorization and normalization. The cache key must include the tenant scope, experiment definition version, time bounds, step, metric, and any filtering dimensions. Otherwise, a fast cache can become a cross-tenant disclosure mechanism or return a cohort definition that no longer matches the label on screen.

Compare operating boundaries, not screenshots

Hosted services differ in query language, retention controls, ingestion accounting, and the work left with the platform team. Their public documentation changes, so this table is a due-diligence map rather than a ranking.

Option Query and interoperability boundary Cost-attribution question to verify Team consequence
Datadog Its Metrics API exposes time-series querying within its platform model. Pricing separates indexed custom metrics and other telemetry categories; confirm which experiment labels create billable series. Less infrastructure ownership, with provider-specific query and billing semantics to model.
Grafana Cloud Hosted metrics documentation describes a Prometheus-compatible service and query path. Confirm active-series, data-point, retention, and plan boundaries against the proposed label budget. PromQL skills can transfer, while account limits and managed-service behavior still require testing.
New Relic Metrics can be queried through NRQL within its telemetry data platform. Map ingest and retained/queryable data terms to each cohort's expected volume. One query model can cover several telemetry types, but the adapter remains specific to that model.

None of those boundaries determines the winner. The evidence should. Run the same representative workload through every candidate: steady card refreshes, a cold 24-hour chart, a bounded 30-day comparison, missing buckets, late samples, and an unauthorized cohort. Record p50 and p95 query latency, error rate, returned-point count, cache hit rate, and ingested series count. Treat those as measurements to collect, never numbers to assume from marketing pages.

The buy-versus-build decision needs the same discipline.

Decision Hosted service Self-hosted metrics stack
Capacity Provider limits and billing dimensions must fit the series budget. The team sizes ingest, storage, compaction, query, and redundancy.
SLO ownership The team owns its adapter and the dashboard; the provider owns a documented service boundary. The team owns the full availability and recovery path.
Lock-in Query syntax, identity, billing, and retention may be provider-specific. Operational procedures and internal expertise become the switching cost.
Cost attribution Bills may expose usage categories, but internal cohort allocation still needs labels or a ledger. Infrastructure cost is visible, yet allocating shared capacity remains an accounting design problem.

I would reject any candidate whose invoice dimensions cannot be reconciled with the cardinality worksheet. That is an explicit trade-off: a slightly nicer query interface is not worth an unexplained growth curve or another after-hours system for a small platform team to own.

Implement a bounded adapter and failure contract

Keep provider syntax inside one adapter. The handler resolves the caller's tenant permissions, loads the immutable experiment definition, validates the request, selects a step that cannot exceed the response-point budget, and calls the adapter with a deadline. It then normalizes series ordering and missing values before returning JSON to React.

Do not retry every failure. Retry a transient upstream timeout only within the request's total deadline and with bounded backoff; do not retry invalid queries or authorization failures. For a stale-but-usable card, the API can return the last successful result with an explicit asOf timestamp and stale state. For experiment guardrails, stale data may be unsafe, so fail closed and make the state visible. The product owner, not the metrics vendor, decides that policy.

Define two SLOs because one hides too much: query availability for valid, authorized requests, and freshness for each decision-critical instrument. A successful HTTP response carrying old samples satisfies the first and violates the second. Alert on sustained budget burn rather than every isolated miss, and keep the dashboard out of its own critical diagnostic path.

Deployment should start with a shadow adapter. Send a sampled set of normalized, read-only queries to the candidate without placing its values on screen; compare timestamps, bucket boundaries, null handling, and aggregates against the existing source. Avoid comparing only totals. Two systems can agree on a daily sum while assigning samples to different cohort windows.

Verify attribution, then make rollback boring

Before release, create fixtures for one authorized tenant in each cohort, one tenant with no assignment, a reassignment exactly at a window boundary, late data, a counter reset, and a time range crossing daylight-saving changes. Store timestamps in UTC and render locale only at the presentation edge. The dashboard should show the unit, interval, cohort-definition version, and freshness without asking the reader to infer them from a tooltip.

The verification gate is concrete: authorization tests pass; the response never exceeds its point budget; observed series stay within the planned label budget; cohort totals reconcile with the assignment ledger; freshness meets its SLO; and measured provider usage maps to an explainable cost category. A single green chart is weak evidence.

Rollback should switch the adapter or hide the experiment card, not require a React deployment. Preserve the old read path during the comparison window, make schema changes additive, and avoid deleting assignment history when an experiment ends. If the new source violates freshness or attribution checks, route reads back, stop shadow traffic, and retain the mismatched query inputs for analysis without retaining credentials or unauthorized tenant data.

The final choice is deliberately unglamorous: select the operating boundary whose measured behavior meets the dashboard SLO, whose label budget survives the next capacity interval, and whose usage can be attributed to tenant cohorts. The API is only the last mile.

References

Top comments (0)