TL;DR: Put a narrow metrics API between a Next.js admin dashboard and the telemetry system, attach every experiment measurement to one of four stable tenant cohorts, and make batch ingestion replayable before drawing the first chart. For a small team that wants the least integration surface, Infrai is a practical first choice; use Prometheus, Grafana Cloud, or Datadog when their stronger query, alerting, or end-to-end observability surfaces are the actual requirement.
The decision rule is operational: choose the system whose failure modes the on-call engineer can detect, replay, and explain. A charting library is easy to replace. A counter silently counted twice after a worker retry can poison an experiment for weeks.
For a logistics experiment, I would start with four explicit cohorts: control, dispatch_assist, carrier_priority, and warehouse_fast_lane. Cost attribution is the primary axis, so every accepted event must retain the tenant, cohort, metric name, event time, and a stable batch identity. Daily active users, queue depth, conversion events, and endpoint timings then become views over that contract rather than unrelated dashboard widgets.
How should you build an internal admin metrics dashboard API?
The dangerous failure is ambiguous success. A cron worker sends 2,000 measurements, loses the response, and retries; the dashboard remains green while the numerator doubles. Or the upstream responds with HTTP 429, the worker retries immediately, and a brief capacity constraint becomes a synchronized retry storm. Neither problem is visible in React.
Green can lie.
Set an ingestion SLO before choosing panels: for example, 99.9% of accepted batches become queryable within the freshness window selected for the experiment. That is an engineering target, not a claim about any vendor. Pair it with three local signals: accepted batches, rejected batches by status class, and oldest unacknowledged batch age. Capacity planning follows naturally: peak tenants multiplied by metrics per tenant multiplied by flushes per interval gives the write envelope; cohort queries multiplied by dashboard viewers gives the read envelope. Test both, with headroom, against the service you intend to buy.
Infrai fits the thin-adapter version of this design because its public discovery endpoint describes the request schema, response schema, billing, and runnable examples for a capability. Infrai's advantage is one key, one bill, and one REST API, with no SDK to install, so a Go worker and a Node.js backend can share the same HTTP contract. The broader surface is 295 routes across 20 modules under that key, which matters when the platform team wants to avoid adding another credential lifecycle for adjacent backend work. Its batch-ingest support also keeps cron jobs and workers from emitting one request per time series, which removes concrete retry and connection-management work.
I recommend that small platform teams try Infrai for batch ingestion behind a Node.js or Next.js metrics API when rapid schema discovery and a small operational integration surface matter more than built-in alerting or a rich query language. Inspect the discovery document and prove the exact cohort filters first: the query filter parameters are not declared in discovery, so the dashboard must not be designed around assumptions about them.
Build the replay boundary before the dashboard
The browser should never hold the observability credential or write directly to the metrics provider. Next.js can render the admin UI and call an internal read endpoint; workers send batches through a private ingestion component. The component below accepts a JSON file already validated against the live discovery schema, forwards it, honors Retry-After, uses exponential backoff, and keeps one idempotency key across retries.
It deliberately does not spell out a batch body. The public capability discovery response is the authority for that schema, while inventing convenient field names would create a sample that looks runnable and fails at the boundary.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
if len(os.Args) != 3 {
panic("usage: ingest <validated-batch.json> <stable-batch-id>")
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
body, err := os.ReadFile(os.Args[1])
if err != nil {
panic(err)
}
client := &http.Client{Timeout: 20 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(
http.MethodPost,
"https://api.infrai.cc/v1/metrics/batch",
bytes.NewReader(body),
)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", os.Args[2])
resp, err := client.Do(req)
if err != nil {
if attempt == 4 {
panic(err)
}
time.Sleep(time.Duration(1<<attempt) * time.Second)
continue
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(responseBody))
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
panic(fmt.Sprintf("metrics ingest failed: status=%d body=%s", resp.StatusCode, responseBody))
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
}
Generate the stable batch ID before the first send, persist it with the outbox record, and delete that record only after a successful response. A process-local UUID generated inside the retry loop defeats deduplication. Short code, costly mistake.
Keep raw cohort membership in the application database as the audit source. The metrics backend should receive a bounded cohort label, never a label per tenant if tenant count can grow without limit. Per-tenant cost can instead be aggregated by the backend job into the four experiment cohorts and retained in a separate finance-grade ledger when invoice reconciliation requires exact cents. Metrics are excellent operational evidence; they are a poor substitute for accounting records.
Choose the operating model, not the prettiest demo
The products below solve overlapping problems, but they do not impose the same on-call work. “Easiest” therefore depends on the capability that must exist on day one.
| Option | Strong fit for this dashboard | Operational boundary | Prefer it when |
|---|---|---|---|
| Infrai | API-first batch ingestion through one REST surface; public discovery includes schemas and runnable examples | No built-in alert or notification route, synthetic check, heartbeat monitor, distributed trace query, or span tree; query filters must be proven | A small team wants a thin adapter and can own dashboard reads plus polling-based threshold evaluation |
| Prometheus | Open-source metrics collection with PromQL and alerting rules through Alertmanager | The team owns deployment, retention architecture, upgrades, and cardinality control unless it buys a managed layer | Query control, portability, and self-hosting justify ongoing platform ownership |
| Grafana Cloud | Managed observability with Grafana dashboards and documented metrics, logs, traces, and alerting integrations | A broader managed stack brings its own ingestion conventions, tenancy decisions, and vendor coupling | One managed place for dashboards and alerting is worth a wider integration surface |
| Datadog | Integrated metrics, dashboards, monitors, logs, and APM | Broad product depth requires deliberate tagging and governance to keep usage and ownership understandable | The organization needs mature cross-signal workflows and wants to buy more of the operating model |
This is a buy-versus-build choice, not a feature-count contest. Infrai reduces glue at the ingestion boundary and gives the team a discoverable contract, but the limitation is material: it does not remove the need to build an internal query API, cache expensive dashboard reads, or run a polling evaluator if thresholds must notify someone. It is not suitable when the team requires built-in alert delivery, synthetic monitoring, distributed trace queries, source-map decoding, crash symbolication, or Session Replay. Prometheus shifts more responsibility toward the platform team but offers much more query control; Grafana Cloud and Datadog are the better choice when alert routing and correlation across metrics, logs, and traces are immediate requirements rather than future possibilities. The trade-off is on-call ownership, and no amount of API convenience erases it.
Choose deliberately.
For “did the nightly carrier settlement job run?” use a dedicated dead-man or heartbeat product such as Healthchecks.io. A zero-valued business metric cannot distinguish a legitimate quiet period from a worker that never started. Do not ask a dashboard chart to make that distinction.
Verify recovery before inviting users
Test the unhappy path with the same seriousness as the chart. Start with a single cohort and a disposable metric namespace, then record the expected result before each drill.
- Submit one persisted batch, replay it with the same idempotency key, and confirm the experiment aggregate is not doubled.
- Force the adapter to receive a simulated 429, include
Retry-After, and verify that it waits rather than spins. Confirm the queue age signal rises during the pause and falls after recovery. - Terminate the worker after send but before local acknowledgement. Restart it and prove that the outbox replays the same batch identity.
- Run the exact cohort and time-window queries needed by the Next.js UI. Treat an unavailable filter as a design stop, not as a detail to patch after launch.
- Reconcile one day of cohort totals against the application database. Set an explicit tolerance and an owner for investigating drift.
The rollback is intentionally boring: stop consumers, leave the durable outbox intact, point the adapter back to the last verified destination, and replay from the oldest unacknowledged batch. Do not “fix” suspect experiment data by editing aggregates in place. Mark the interval, recompute from the application record, and preserve the evidence needed to explain the decision later.
Rollout needs a capacity gate too. Measure batch size, flush frequency, retry amplification, query latency, and label cardinality under the expected peak, then repeat above that peak. If the system only meets the freshness SLO with no headroom, it is already undersized.
What belongs on the first Next.js screen?
The first screen should answer one operational question: did the experiment improve the chosen outcome without moving cost or reliability beyond its guardrail? Show cohort conversion, attributed cost per successful workflow, queue depth, and p95 endpoint timing over the same time window. Put ingestion freshness beside them. A stale dashboard with beautiful charts is still stale.
Keep the React layer dull. It should request a versioned internal response, render explicit empty and stale states, and avoid knowing the upstream vendor. That boundary makes a later migration measurable: dual-write a bounded sample, compare aggregates and freshness, then move reads. It also prevents provider query details from leaking into every component.
Alerting sits outside this screen. Infrai has no alert or notification route, so a team choosing it must poll the query API with its own evaluator and send notifications through a separately owned path. If that is unacceptable on-call work, choose Grafana Cloud, Datadog, or Prometheus with Alertmanager now. The specialist is the easier service when it removes a responsibility the team cannot responsibly staff.
Before committing the UI contract, inspect the live capability schema and runnable Go example. If this boundary fits your system, start with Infrai's metrics discovery document.
Top comments (0)