The operational constraint for a startup SaaS team choosing Grafana Cloud or a simple custom metrics API is attribution: if a Node.js AI agent loop can retry, call several models, and invoke tools for one tenant, the easiest dashboard embed is not yet a trustworthy business metric. Choose the storage and presentation path only after one immutable usage record can explain who caused each unit of work across US and EU workloads.
TL;DR: define an append-only cost ledger at the agent boundary, keep display labels out of metric identity, and make every dashboard total reconcilable against that ledger. A hosted observability stack and a small metrics API are delivery options, not competing definitions of truth. The right choice is the one that preserves tenant isolation, US/EU placement, bounded cardinality, and an auditable path from a chart cell back to recorded work.
Consider a bounded incident exercise for a B2B SaaS agent: a customer sees its daily AI usage jump after an agent retries one tool call. The dashboard groups work by region and model, but it cannot distinguish a customer request from an internal retry. Support cannot explain the invoice preview; engineering cannot tell whether the graph double-counted the attempt or the retry incurred more work. I would classify that as an attribution incident even if every service stayed available, because the number has lost the property the business needs: explainability.
The invariant is blunt. Every chargeable attempt needs one stable owner, one stable operation identity, and one terminal accounting outcome. Aggregation comes later.
No shortcut fixes that.
Why did the dashboard total become impossible to defend?
Most teams begin with the visible question: how quickly can we embed a chart? That reverses the dependency. The chart can only repeat the semantics of the events beneath it, and an agent loop has awkward semantics: one user action can fan out, a retry can be legitimate, a tool can fail after consuming capacity, and a response can complete after the browser has gone away.
A request counter tagged with tenant_id, user_id, conversation_id, run_id, tool_name, and free-form error text looks flexible. It also mixes three jobs. It is trying to be an operational signal, an audit record, and a billing ledger. Those jobs have different retention, query, and identity requirements. One metric series should not carry every business dimension merely because the dashboard wants a filter.
The specific failure is usually deduplication. Suppose attempt 1 times out from the caller's perspective and attempt 2 succeeds. A counter emitted only at the HTTP edge can count two requests without describing which attempt consumed model or tool work. A counter emitted only at successful completion can hide consumed work from attempt 1. Neither number is necessarily false; each answers a different question. Calling both agent_requests_total turns ambiguity into policy by accident.
Three words matter: name the denominator.
For an SLO, I may need successful user-visible runs divided by eligible runs. For capacity, I need attempts and their consumed units. For customer cost attribution, I need the accounting policy's accepted usage records, including an explicit treatment for failed or retried work. Those are separate projections of the same execution, and the dashboard should say which projection it shows.
Should a startup SaaS use cloud metrics or a custom API?
I would record a compact event where the agent orchestration layer knows both the tenant and the attempt. The event needs a deterministic identity that survives retry handling, plus dimensions that are governed rather than copied blindly from request metadata. Keep the example deliberately small: tenant, region, operation, attempt, outcome, and measured usage.
Grafana Cloud is the named managed option in this choice, while a custom metrics API describes a system boundary rather than a comparable product. That mismatch matters. The managed path can be evaluated as an operating model; the custom path still requires decisions about storage, queries, authorization, backup, regional routing, and on-call ownership. Treating the latter as “just an API” hides most of the work. Treating the former as “just a dashboard” hides the instrumentation and accounting policy that remain with the SaaS team. Neither label answers whether a replay changes a tenant total, so neither can be selected from the embed demonstration alone.
The following Go type is a boundary contract, not a vendor SDK. Integer usage fields avoid rounding during aggregation; OccurredAt supports placement into an accounting window; EventID is the idempotency key used by the receiving store. The exact units and acceptance policy belong in a versioned internal specification.
package usage
import (
"context"
"errors"
"time"
)
type Event struct {
EventID, TenantID, Region, Operation, Outcome string
Attempt int
InputUnit, OutputUnit, ToolCalls int64
OccurredAt time.Time
}
type Recorder interface {
Record(context.Context, Event) error
}
func RecordTerminal(ctx context.Context, r Recorder, e Event) error {
if e.EventID == "" || e.TenantID == "" || e.Attempt < 1 {
return errors.New("invalid usage event identity")
}
if e.Region != "us" && e.Region != "eu" {
return errors.New("unsupported accounting region")
}
switch e.Outcome {
case "succeeded", "failed", "cancelled":
default:
return errors.New("invalid terminal outcome")
}
if e.InputUnit < 0 || e.OutputUnit < 0 || e.ToolCalls < 0 {
return errors.New("usage cannot be negative")
}
return r.Record(ctx, e)
}
This code does not decide whether a failed attempt is chargeable. That is intentional. It makes the raw terminal outcome available so policy can be applied explicitly and tested. I would add a schema version and validate operation names against a controlled registry; I would not add user email, prompt text, or arbitrary exception messages to the accounting key.
Policy stays visible.
Emission needs a failure policy. Blocking the customer response may protect accounting completeness but damage availability. Buffering and retrying can preserve the user path, but it creates backlog capacity and recovery objectives that the on-call team must own. There is no universal answer. The decision should state an error budget for missing or late usage events, a maximum acceptable ledger lag, and what happens when that budget is exhausted.
Compare operating models on reconciliation, not screenshots
A managed observability service, a self-hosted stack, and a purpose-built metrics API can all display totals. I would take each option through the same buy-versus-build review rather than score ease from a five-minute demo. No option gets credit for capabilities the team has not tested against its identity and residency boundaries.
| Decision test | Managed observability | Self-hosted observability | Purpose-built metrics API |
|---|---|---|---|
| Accounting truth | Keep an independent ledger; verify export and replay | Keep an independent ledger; own storage operations | Do not let an aggregate endpoint become the only record |
| On-call load | Team still owns instrumentation and policy | Team owns upgrades, capacity, backup, and recovery | Team owns API correctness, query behavior, and storage |
| Cost attribution | Validate effects of dimensions, retention, and queries | Plan series, event volume, retention, and replicas | Plan writes, deduplication state, retention, and read load |
| Lock-in test | Can accepted events and definitions be exported and replayed? | Can components change without changing event meaning? | Can another implementation pass the contract tests? |
| US/EU boundary | Verify placement and support access by data class | Operate isolated regional paths and prove backup placement | Route regional writes and reject ambiguous placement |
| Embedded access | Test authorization at tenant scope | Build and operate the authorization boundary | Make tenant-scoped querying an API invariant |
This table avoids a fake winner. Managed service costs include more than a bill: instrumentation changes, incident coordination, and exit work remain yours. Self-hosting has no licensing shortcut around on-call labor. A narrow API can reduce surface area while concentrating responsibility for correctness in code the team owns.
The limitation of the managed path is that it cannot define the team's accounting semantics; the limitation of the custom API is the operational surface the team must build and support. Self-hosting is not suitable when the team cannot staff backup recovery, upgrades, and capacity incidents. A custom API is not suitable when tenant-scoped authorization and idempotent storage are still assumptions. Managed observability is not sufficient when accepted raw usage cannot be replayed and reconciled under the team's retention and regional rules. These are boundaries, not rankings.
Capacity planning should use workloads, not adjectives. Estimate terminal events per agent run, peak runs per second, retry amplification, bytes per accepted event, retention windows, regional duplication, and the largest reconciliation query. Then load-test the ugly case: a retry storm during a regional backlog while finance requests a tenant-level daily total. Average traffic is not the sizing input that wakes someone up.
For a concrete fixture, one successful attempt with 800 input units and 120 output units should still total 920 after the identical event is submitted twice. A second event with a distinct attempt identity must remain distinct even if it belongs to the same run. That tiny dataset tests the important trade-off: deduplicate replays without collapsing legitimate retries. It also makes a failed review actionable. If a backend returns 1,840 units, its acceptance boundary did not enforce idempotency; if it returns no underlying event IDs, the aggregate cannot be independently reconciled; if a US reader can silently query the EU fixture, regional routing or authorization is incomplete. The numbers are examples for the contract below, not a benchmark or a pricing claim.
I would demand one proof from every candidate: ingest the same fixture twice, query it through the intended tenant boundary, export the accepted records, and reconcile the result by event identity. If the total changes after replay, the design does not yet have idempotency. If an operator can query another tenant without an audited support path, the embed is not ready.
The preventative path is a contract test
The useful test crosses the ledger and the aggregate. It records a successful first attempt, replays that event, and checks the policy projection. Dashboard tests alone cannot establish this because a plausible pixel is not evidence of correct accounting.
package usage_test
import (
"context"
"testing"
"time"
"example.invalid/usage"
)
func TestReplayDoesNotChangeAcceptedUsage(t *testing.T) {
ctx := context.Background()
store := newMemoryStore()
at := time.Date(2026, time.October, 4, 12, 0, 0, 0, time.UTC)
first := usage.Event{
EventID: "run-42-attempt-1", TenantID: "tenant-a",
Region: "eu", Operation: "agent_reply", Attempt: 1,
Outcome: "succeeded", InputUnit: 800, OutputUnit: 120,
ToolCalls: 2, OccurredAt: at,
}
if err := usage.RecordTerminal(ctx, store, first); err != nil {
t.Fatal(err)
}
if err := usage.RecordTerminal(ctx, store, first); err != nil {
t.Fatal(err)
}
if got := store.AcceptedUnits("tenant-a", at); got != 920 {
t.Fatalf("accepted units = %d, want 920", got)
}
}
The store implementation is omitted because its required behavior is the point. The production store must enforce uniqueness atomically at the boundary where it acknowledges acceptance. A read-before-write check without that guarantee can still double-count under concurrency.
Operational logs should explain why a record was rejected or delayed without becoming a second cost ledger. If severity labels cross systems, map them deliberately. RFC 5424 defines numerical severity values and states that each message has one severity; it also notes that the meaning of locally used levels can differ. That is a reason to document mappings, not to infer accounting outcomes from a log label.
When this design is too much
An append-only attribution ledger is unnecessary for a dashboard with no tenant-level money, quota, entitlement, or contractual decision attached to it. A small internal operational panel can start with bounded metrics, provided everyone accepts that its totals are approximate and cannot later serve as invoice evidence.
It is also the wrong abstraction for debugging payloads. Traces and logs answer execution questions with context that should not be copied into a cost record. Link them through a controlled run identifier, apply separate access and retention rules, and keep sensitive content out of business dimensions.
The decision changes when attribution becomes consequential. Choose the dashboard backend after the ledger passes replay, isolation, regional placement, and reconciliation tests. This ordering leaves room to buy, host, or build the presentation layer later without rewriting what a unit of agent work means. It also gives the SLO review a real question: not merely whether charts loaded, but whether accepted usage remained complete, timely, and explainable under failure.
Top comments (0)