A customer-support metrics system has one constraint that changes the answer: a dashboard may be wrong briefly, but it must not destroy the evidence needed to reconstruct a customer incident. For a startup MVP, choose a managed analytics layer over a custom ingestion API when the source database can retain immutable evidence for 30 days and every derived change can be rolled back independently.
TL;DR: keep raw support events append-only, derive replaceable metrics from them, and buy the dashboard layer first. Build the API when the write contract is itself part of the product, or when event volume, isolation, or retention requirements make direct analytical access an unacceptable boundary.
What must survive a bad deployment?
A support lead may need to establish when a conversation entered a queue, which policy version classified it, when an agent responded, and whether a correction changed the reported result. A daily count cannot answer that. Neither can a mutable row containing only the latest state.
The durable record needs stable event, tenant, and conversation identifiers; event and ingestion times; an event type; a schema version; and enough payload to reconstruct the action. Corrections should be new events referring to earlier ones. If a parser is wrong, deploy a corrected projection and replay the retained interval without rewriting history.
Thirty days is a design boundary here, not a universal recommendation. Pick another window if support policy demands it, then test storage capacity and recovery time. Dashboard retention is not evidence retention: an aggregate may outlive its inputs while remaining impossible to audit.
Rollback has two meanings. Code rollback restores a projector. Data rollback restores the previous interpretation of accepted events, which requires versioned definitions, reproducible replay, and a way to switch readers between old and new results.
Preserve the input.
Derive the architecture from the constraint
The evidence plane commits support events durably before acknowledging success. The presentation plane reads derived tables or views and may be rebuilt. A queue can separate them, but it is not mandatory if the database transaction provides the required acceptance semantics and projection work stays off the request path.
A metric row should identify its definition version and covered interval. Compute version 2 beside version 1, validate both, move readers, and retain version 1 until the rollback window closes. Late events also need an explicit policy: recompute closed intervals, mark them provisional, or reject lateness beyond a stated limit.
from dataclasses import dataclass
from datetime import datetime
from typing import Protocol
@dataclass(frozen=True)
class SupportEvent:
event_id: str
conversation_id: str
event_type: str
occurred_at: datetime
schema_version: int
class EvidenceStore(Protocol):
def append_once(self, event: SupportEvent) -> bool:
"""Commit once; return False if event_id already exists."""
def accept(event: SupportEvent, store: EvidenceStore) -> dict[str, str]:
inserted = store.append_once(event)
return {"event_id": event.event_id,
"status": "accepted" if inserted else "duplicate"}
This contract does not prove end-to-end exactly-once processing or order events across producers. Its useful property is narrower: retrying a stable identifier cannot create a second evidence record, and projections remain traceable to retained inputs.
Prometheus instrumentation guidance recommends labels rather than generated metric names, while warning that each unique label combination creates another time series and that high-cardinality values such as user IDs or email addresses should not be labels. A bounded queue or outcome may fit; conversation_id, message text, and email belong in the evidence store.
Should you build or buy a metrics dashboard backend?
Treat the shortlist as deployment shapes, not endorsements. “Supabase/Postgres charts” means the team owns event tables, projections, access boundaries, and chart interfaces. “Metabase” means a database-connected analytics layer reads controlled derived relations. “Grafana” means a metrics-oriented dashboard reads bounded time-series dimensions while evidence stays elsewhere. These are comparison boundaries, not claims that a product guarantees an architecture.
| Path | Rollback boundary | Failure mode to test | Choose only if |
|---|---|---|---|
| Supabase/Postgres charts | Team versions ingestion, projection, and UI | Schema changes couple evidence writes to releases | Interface control justifies owning replay and isolation |
| Metabase-style analytics | Derived relations roll back independently | Ad hoc reads reach evidence or contend with capture | Read-only roles and workload limits are enforced |
| Grafana-style dashboards | Evidence replay and metric publishing stay separate | Cardinality expands through customer identifiers | Every label has a finite budget |
| Separate ingestion API | API schema and evidence log form the boundary | Retries or partial writes create gaps | The team can operate durability, replay, quotas, and migrations |
My choice for this MVP is managed analytics over curated Postgres relations and an append-only evidence table. The hard problem is reconstructability; a bespoke chart API creates another contract before the event model is stable. Safety comes from permissions, isolation, versioned projections, and rehearsed replay, not a product name.
The custom path becomes preferable when capture needs independent scaling, external clients need a stable ingestion contract, or analytical reads cannot be isolated from writes. Then the API controls failure domains rather than merely replacing SQL with HTTP.
This recommendation has a hard limitation: managed analytics is not suitable when dashboard readers cannot be isolated from the evidence writer, when external producers require a stable public contract, or when replay cannot finish inside the declared recovery objective. In those cases, choose the separate ingestion API and accept its operational burden. The reverse trade-off matters too. A custom API is not suitable merely because the team prefers application code to SQL; without duplicate handling, backpressure, schema compatibility, durable acceptance, and replay ownership, it adds a failure boundary without buying rollback safety. The deciding evidence should come from a load test and a replay drill using the team's stated limits, not from a feature checklist.
Failure modes worth testing
A green chart is weak evidence. Send the same event twice. Deliver events out of order. Introduce the next schema version while the current projector runs. Publish a faulty definition, recompute seven days beside its predecessor, switch readers, and switch back. Exhaust the label budget. Run an expensive analytical query while support traffic continues.
Retries are normal.
Silent loss isn't.
Write down the maximum event size, request deadline, replay window, recovery objective, allowed lateness, query timeout, and cardinality ceiling before testing. No universal number can be inferred for this system.
If an event commits but its response is lost, the producer retries; stable IDs make the retry harmless. If projection fails after acceptance, evidence remains and lag alerts separately. If a derived migration fails halfway, readers still point to the last complete version. If an external notification must follow a commit, record an outbox entry in the same transaction and deliver it asynchronously.
Cost belongs in the review, but price is not the architecture. Model evidence bytes, index amplification, projection compute, dashboard concurrency, cardinality, replay frequency, and operator time. Commercial observability pricing may separate ingestion from indexed or retained data; measure both flows instead of selecting from a transient unit price.
Roll out without wagering the ledger
Start with one support workflow and a shadow projection. Keep the existing report authoritative while the new path consumes the same bounded event set, then compare totals and sampled reconstructions across the 30-day window. Record discrepancies by event ID, not screenshot.
Expose only curated, read-only relations to the dashboard identity. Apply query timeouts and resource limits, alert on projection lag and rejected events, and rehearse switching readers to the previous definition without deleting either version.
Choose managed analytics while immutable inputs, isolated reads, and versioned projections make the dashboard replaceable; choose a custom ingestion boundary when capture requires independent scaling or a stable external contract. In both cases, preserve enough evidence to explain the chart after the chart is wrong.
Top comments (0)