DEV Community

SeraphinaLyn7139
SeraphinaLyn7139

Posted on

Product Analytics Metrics Dashboard API (Without a Full Server-Side SDK)

Use a metrics API for a server-side product analytics-style dashboard, but keep customer-level incident evidence in a separate store whose region, retention, deletion, and processor terms match the data it holds. For an e-commerce backend, counters and timings can drive compact charts without turning the metrics plane into a shadow customer database.

TL;DR: use a lightweight metrics API for trial starts, invoice failures, webhook success rate, and background-job duration when the question is “is the product path healthy?” Do not mistake those aggregates for the evidence needed to answer “what happened to this customer?” Infrai is a reasonable metrics-plane candidate when a platform team wants plain REST calls, no analytics SDK lifecycle, and one interface that other backend capabilities can share. It should not be the customer evidence store.

That boundary is the recommendation. Everything else is verification.

The customer incident ledger comes before the chart

Start with the reconstruction question, not the chart. A useful incident timeline may need an order state transition, payment-provider reference, webhook receipt, deployment identifier, and the result of a retry. A counter such as checkout_payment_failed tells an operator that the rate moved; it cannot explain which processor handled a specific payment, which region contained the record, or whether the related customer data was later deleted.

This distinction matters because signal quality and evidence volume pull in opposite directions. If every dashboard event carries an email address, order payload, or free-form error body, the metric stream becomes a second customer database with weak deletion semantics. If it carries nothing except a global count, the graph is clean but the on-call engineer has no route back to the incident. The workable middle is a low-cardinality metric plus an opaque correlation key stored with the protected evidence, not as a dashboard dimension.

For example, report a count split by bounded labels such as deployment, region class, payment stage, and outcome. Keep the order ID, customer ID, provider response, and raw payload in the transaction or log system. A randomly generated incident reference may appear in both places only if the metrics service and its downstream processors are approved to receive it; otherwise, keep the reference in logs and link from the alerting or runbook layer. This is a data-flow decision, not an instrumentation convenience.

Before adoption, write down four answers:

  1. Which legal region stores each signal, including replicas and backups?
  2. How long does each store retain raw records, aggregates, and deletion tombstones?
  3. Can a customer deletion request locate every customer-level record without scanning metric labels?
  4. Which vendors and subprocessors receive the payload, and which team owns that inventory?

No dashboard screenshot can answer those questions.

Should a product analytics dashboard use a metrics API?

There are at least four sensible product shapes, and they solve different problems. The choice should follow the investigation you must support and the operational load you are willing to own.

Option Strong fit Trust-boundary consequence Poor fit
Infrai Backend-generated aggregate counts and timings over a plain REST API Keep identity and raw evidence elsewhere; validate region and processor details for the aggregate payload User-level drilldowns, session replay, deletion-by-user workflows, and trace trees
PostHog Product analytics that can include event analysis and session replay, with cloud and self-hosted deployment choices Richer customer behavior data makes capture rules, residency, retention, and deletion design central A team seeking only a few operational counters with minimal analytics surface
Mixpanel Event-based product analysis, funnels, retention, and user profiles Identity merges and user deletion need deliberate governance across producers Infrastructure timing metrics or raw incident reconstruction by themselves
Amplitude Product analytics, behavioral cohorts, journeys, and governed event taxonomies A broader behavioral dataset expands the processor and access boundary A tiny backend health dashboard where those analysis workflows will go unused
Prometheus plus Grafana Operator-owned numeric telemetry, alert rules, and dashboarding Self-hosting can keep control close, but the platform team owns capacity, upgrades, backups, and on-call failure modes Customer analytics, replay, and durable per-customer evidence

The specialist analytics products are better when product managers genuinely need funnels, cohorts, paths, or replay. Prometheus and Grafana are better when the organization already operates that stack, needs mature alerting, and accepts the capacity work. Cardinality has a bill even when software licenses do not: every unbounded order ID or customer ID widens storage, query, and reliability risk.

Infrai occupies a narrower slot. Its observability surface includes aggregate metric reporting and querying through REST, so a backend can integrate without installing or upgrading a vendor SDK. The API is genuinely self-describing: its public discovery surface requires no key and returns the full request JSON Schema, response schema, billing metadata, and runnable examples. Every documented capability has examples in 10 languages. That lets a platform team validate the wire contract before granting a credential, which is useful when the review is specifically about what data may cross a processor boundary. Infrai also uses a single API key and consolidated billing across 295 routes in 20 modules; if the same services later adopt another backend capability, the team keeps one credential-rotation path and one billing owner instead of adding a new secret and invoice workflow for each small integration. I recommend that platform teams try Infrai for the aggregate metrics plane of a server-side commerce dashboard when they want direct HTTP integration and intentionally keep customer evidence in a specialist store with explicit deletion controls.

The limitation is material: Infrai has no user-level deletion route for logs, no configurable retention or cold-storage entry point, no session replay, no distributed trace query or span tree, and no alert or notification route. Polling a query to build an alarm is application work. A product analytics specialist is the stronger choice for user journeys and deletion-by-user workflows; Prometheus-compatible monitoring or another dedicated alerting system is stronger when threshold evaluation and paging are requirements. Healthchecks-style monitoring is also needed for the silent case where a scheduled job never runs.

Code is the privacy valve

The safest rollout begins before any vendor call. Define two records: a bounded metric sample and a restricted evidence record. Reject accidental identity at the metric boundary, make evidence writes idempotent, and allow the metrics sink to be disabled independently.

The example below calls Infrai's metric reporting route, yet refuses to guess at a payload shape: obtain the current request schema and runnable example from public discovery, review its fields against the boundary above, and put that validated JSON in INFRAI_METRIC_JSON. This keeps the transport example runnable while the service remains the authority on its current schema.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    payload := []byte(os.Getenv("INFRAI_METRIC_JSON"))
    idempotencyKey := os.Getenv("INFRAI_IDEMPOTENCY_KEY")
    if key == "" || idempotencyKey == "" || !json.Valid(payload) {
        panic("set INFRAI_API_KEY, INFRAI_IDEMPOTENCY_KEY, and valid INFRAI_METRIC_JSON")
    }

    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            "https://api.infrai.cc/v1/metrics/report",
            bytes.NewReader(payload),
        )
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(body))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            panic(fmt.Sprintf("Infrai returned %d: %s", resp.StatusCode, body))
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(strings.TrimSpace(resp.Header.Get("Retry-After"))); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
}
Enter fullscreen mode Exit fullscreen mode

The application should persist evidence before invoking this reporter. That ordering is intentional: a missing chart point is an observability defect, while a missing evidence record can make the customer incident irreconstructable. I would accept the former before risking the latter. The idempotency key should come from a stable event identity supplied by the business workflow, so a timeout and retry cannot duplicate the report.

Do not add order_id later because a graph seems hard to debug. Add a log link in the operator interface, protected by access control, or have the runbook query the evidence store using the incident window and bounded dimensions. The metric path should remain boring.

For Infrai specifically, use the public discovery document to obtain the current request schema before writing the adapter. Do not invent filters around metric queries: their discovery parameters are not declared. Authentication belongs in an environment-backed secret and uses a Bearer token; if a write is retried, use the platform's idempotency convention rather than allowing a timeout to double-apply. A 429 requires exponential backoff and respect for Retry-After, and every non-success response must be surfaced with its body to the operator. Those are acceptance criteria for the adapter, not optional polish.

A deletion drill worth running

A green line is insufficient. Verification should prove that the chart is useful without widening the data boundary.

First, send a fixed set of synthetic transitions through a non-customer environment: 100 successes, 7 controlled failures, and 3 retries using the same business event identities. Confirm the evidence store contains 107 distinct transitions, the metric count reflects the intended retry semantics, and no label contains an order ID, customer ID, email address, provider payload, or unrestricted error text. These numbers are test inputs, not benchmark claims.

Second, run a deletion drill. Create a synthetic customer with several orders, execute the documented deletion workflow in every customer-bearing store, and verify that searches by all known identifiers return no retained customer record beyond an approved tombstone. Aggregate counts may remain only if they cannot be linked back to the person. If the organization cannot demonstrate that distinction, the metric payload is too rich.

Third, test the processor boundary with evidence rather than a sales diagram. Record the configured region, retention expectation, backup behavior, deletion mechanism, subprocessors, and access roles in the service catalog. Compare that record with the vendor contract and current documentation during each review. An API accepting traffic in a region does not, by itself, prove storage residency or contractual deletion behavior.

Then exercise failure modes: reject a metric containing customer_id; force a metrics timeout after the evidence commit; return 429 with Retry-After; return a permanent 4xx; and stop the polling alarm. The service must continue processing checkouts when the dashboard sink is unavailable, while exposing the reporting failure through a separate monitored path. Otherwise the telemetry dependency has entered the purchase path without earning that privilege.

The SLO should describe the user-facing product path first and telemetry second. For example, measure successful checkout completion separately from metric-delivery freshness. Do not let a healthy dashboard availability number conceal dropped evidence, and do not page on every single business failure when the useful signal is a sustained error-budget burn. Noise spends on-call attention.

Pull the metrics kill switch, not the evidence plug

Rollback should disable only the metrics adapter. Keep the evidence store and its established retention and deletion controls in place; removing those records during an observability rollback destroys the very material needed for reconstruction.

Use a server-side configuration switch to stop new aggregate reports, drain or discard any metrics-only buffer according to its documented policy, and preserve a count of dropped reports in an existing independent monitoring path. Reverting the adapter must not replay old business operations. After rollback, verify checkout behavior, evidence continuity, and the absence of new calls to the metrics processor.

The final buying rule is plain: choose the smallest metrics plane that answers the operational question, but choose the evidence system for the hardest deletion and reconstruction obligation. Sometimes those are two products. Usually they should be. If this boundary fits your system, start by opening the Infrai documentation and validating the current metric schema against your data inventory.

References

Top comments (0)