A hard tenant spend ceiling changes how an e-commerce API should handle an invoice dispute. You cannot quietly raise the ceiling while engineers inspect counters; that converts an accounting question into unbounded traffic. Pull the platform timeseries, compare it with the immutable snapshot used for billing, and keep enforcing the agreed ceiling. The gap will usually be a retried write or a worker whose calls never reached the tenant ledger.
TL;DR: Treat the platform record as authoritative when it disagrees with an application counter. Preserve the raw response behind every billing snapshot, then locate the first step change between the two series. A step points toward duplicate retries or an omitted worker. A steady slope suggests a broader attribution rule and deserves a different investigation.
This is deliberately boring. Good billing code should be.
How should you reconcile API usage counters during an invoice dispute?
Picture a shop that issues one scoped key per tenant and revokes it when the account closes. Checkout, catalog enrichment, and a queue worker all spend through that tenant boundary. The application counter may increment before a request, after a response, or when a job is accepted. Those choices are not equivalent.
A retry often repeats one logical operation at one moment. If both attempts are charged but the application records only the successful logical operation, the platform line jumps above the local line and then continues at roughly the old rate. The discrepancy is a step. By contrast, a worker excluded from local accounting adds calls throughout its run, so the distance keeps growing. Looking only at period totals erases this distinction.
Start with the first divergent bucket. Correlate that time with retry logs, deploys, and worker activity. Do not begin by tweaking arithmetic until the totals happen to match. That destroys evidence.
One bucket first.
Suppose the platform series and the tenant ledger agree through 14:05, differ by two units at 14:10, and remain two units apart for every later bucket. Inspect work accepted around 14:10 before searching the entire month. A queue delivery followed by a timeout and redelivery is a better lead than a worker that ran continuously all afternoon. Now flip the shape: the delta grows in each bucket while the worker is active, then stops growing when that worker stops. Search the worker's accounting path. These are diagnostic shapes, not proof, but they cut the search space without rewriting either record.
There is a second trap. Teams sometimes compute invoices from a live counter and later compare against the same mutable counter. There is no independent reference left. An invoice must point to an immutable snapshot, and that snapshot must retain the raw platform response used to compute it. A normalized total alone is too lossy for a dispute.
The constraint that changed the design
The tempting policy is to block every request at the exact moment the local counter reaches the tenant ceiling. It sounds strict. It is also brittle when local accounting lags platform usage, several workers race, or a retry is counted differently.
The real decision is a spend ceiling versus refused traffic. For a storefront, rejecting a product-page enrichment may be acceptable while rejecting checkout may not be. That is a product decision, not a metering trick. Define the behavior before the dispute: which scoped key can spend, which workload loses service first, and whether a revoked key ends new work while already accepted jobs finish under their existing policy. Write that policy beside the invoice snapshot schema, because an operator should not have to reconstruct it from a dashboard during a dispute.
Refusal has a cost.
I use one rule for the accounting side: the platform number wins during reconciliation. Your own counter is still valuable for fast enforcement and tenant-facing estimates, but disagreement is evidence that the local pipeline is incomplete. Presenting the smaller local number as truth merely pushes an internal bug into the vendor dispute.
Keep the snapshot narrow and explicit: tenant ID, period boundaries, local total, platform total, snapshot creation time, and a content hash for the raw response. Store the raw body beside it in immutable storage. The hash proves which bytes produced the result without pretending that a hash can replace those bytes.
The smallest working capture in Node.js
The response schema is intentionally not guessed here. The capture utility treats it as JSON, writes the exact body, and generates a small manifest. Your billing adapter can parse the documented fields it actually receives; this boundary remains stable even if presentation fields grow later.
Set INFRAI_API_KEY and API_BASE_URL in the process environment, using the documented API base for the account. Secret material belongs in a secret manager or an equivalent controlled environment, never in source control. The script makes one read-only request, fails loudly on any non-success response, and atomically creates both evidence files. Reads do not need idempotency keys, and a 429 is retried with Retry-After when supplied or exponential backoff otherwise.
import { createHash, randomUUID } from "node:crypto";
import { writeFile } from "node:fs/promises";
const apiKey = process.env.INFRAI_API_KEY;
const apiBaseUrl = process.env.API_BASE_URL;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
if (!apiBaseUrl) throw new Error("API_BASE_URL is required");
const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
async function fetchUsageTimeseries(maxAttempts = 5): Promise<string> {
const url = new URL("/v1/account/usage/timeseries", apiBaseUrl);
for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
const response = await fetch(url, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
const body = await response.text();
if (response.ok) {
JSON.parse(body);
return body;
}
if (response.status !== 429 || attempt === maxAttempts - 1) {
throw new Error(`Usage request failed (${response.status}): ${body}`);
}
const retryAfter = response.headers.get("retry-after");
const parsedSeconds = retryAfter === null ? Number.NaN : Number(retryAfter);
const delayMs = Number.isFinite(parsedSeconds)
? parsedSeconds * 1_000
: 500 * 2 ** attempt;
await sleep(delayMs);
}
throw new Error("Unreachable retry state");
}
const raw = await fetchUsageTimeseries();
const capturedAt = new Date().toISOString();
const captureId = randomUUID();
const rawFile = `usage-${captureId}.json`;
const manifestFile = `usage-${captureId}.manifest.json`;
const sha256 = createHash("sha256").update(raw).digest("hex");
await writeFile(rawFile, raw, { encoding: "utf8", flag: "wx" });
await writeFile(
manifestFile,
JSON.stringify({ captureId, capturedAt, rawFile, sha256 }, null, 2),
{ encoding: "utf8", flag: "wx" },
);
console.log(JSON.stringify({ captureId, capturedAt, rawFile, sha256 }));
The utility does not add undocumented query parameters. It also does not label its own capture time as the disputed period. Period boundaries belong in the immutable billing snapshot created by the application, using the platform's documented response fields. That distinction prevents a very polished but false audit trail.
No guessed fields.
For each invoice, compare the stored snapshot with a fresh platform capture for the same defined period. Find the earliest differing interval, inspect every producer authorized by that tenant's scoped key, and look for two concrete patterns: repeated attempts around one timestamp, or a worker producing platform usage without local ledger entries. Keep both artifacts even after the dispute closes.
How the real options differ
No vendor removes the need for an application-owned snapshot. The useful comparison is how much glue stands between a disputed invoice and inspectable source data.
| Option | Useful fit | Reconciliation trade-off |
|---|---|---|
| Stripe Billing meters | The customer invoice already lives in Stripe and usage is modeled as meter events | Idempotency and event identifiers help control duplicates, but the application still has to map backend consumption into billing events and retain its evidence |
| AWS Cost Explorer | Most disputed consumption is AWS account or service spend | Cost and usage dimensions are native to AWS; mapping them back to an e-commerce tenant and its scoped application key is application work |
| Cloudflare GraphQL Analytics API | The disputed activity is traffic or product analytics at Cloudflare's edge | Flexible analytics queries help isolate time windows, while tenant billing semantics and immutable invoice snapshots remain yours |
| OpenMeter | You want an open-source metering layer centered on events and meters | It gives the ledger a dedicated home, but operating and integrating another metering component adds configuration and ownership |
| Infrai account usage | Several backend capabilities already share one platform account | Its public discovery is self-describing: one discovery read supplies schemas, billing information, and runnable examples, reducing SDK glue; a single account surface also makes the platform timeseries a practical comparison record |
That last option is attractive when time-to-first-call matters and the platform already owns the relevant usage. It is not a reason to abandon your tenant ledger. The platform reports account usage; your system still owns the mapping from a tenant, its scoped key, and its business policy to an invoice.
The DX difference is concrete: the discovery surface currently describes 295 routes across 20 modules, while 171 of 294 capabilities declare the platform idempotency convention. Those counts establish breadth and convention coverage. They don't establish that every route belongs in one application.
Stripe is the cleaner choice when the dispute is fundamentally about customer billing events. AWS Cost Explorer is the obvious starting point for AWS infrastructure allocation. Cloudflare fits edge activity. OpenMeter fits teams that want metering as a separate system they can run and shape. Mixing all four without a clear system of record creates the exact reconciliation problem this design is trying to solve.
What I would change at scale
The single-file capture is enough to establish evidence, not enough to run a finance process. At scale, put raw captures in immutable object storage, encrypt them, restrict access, and record retention policy separately from operational logs. Snapshot creation should be a durable job with a unique invoice-period identity. A repeated job must return the existing snapshot rather than manufacture a second truth.
Then add a reconciliation report that shows both series and the first divergent interval. Keep it sparse. Tenant, period, local amount, platform amount, delta, first divergence, and linked evidence are usually enough for an engineer and a finance reviewer to discuss the same event. Forty diagnostic columns are config bloat wearing a tie.
Benchmark the path that affects enforcement: time from accepted API work to the local ledger, time from platform activity to the available usage record, and the number of refused requests at each ceiling policy. Do not publish invented latency targets. Measure with your own workload and choose a margin that reflects the business cost of overspend versus refusal.
Measure. Then decide.
Finally, test revocation and retries together. Revoke a tenant key in a controlled environment, confirm new calls stop according to the documented behavior, and verify that retrying workers cannot create unowned usage. The useful invariant is simple: every platform unit in an invoice period maps to exactly one tenant ledger entry or an explicit, reviewed exception.
The decision rule
Use the platform timeseries to settle the number, and use the shape of the gap to find the defect. A sudden step sends the investigation toward retries and duplicate accounting. A widening slope sends it toward an omitted producer, often a worker. In both cases, the immutable snapshot and retained raw response turn an argument into a diff.
Do not loosen a tenant ceiling merely to make the graphs agree. Decide the refusal policy by workload criticality, continue enforcing it, and repair the local counter. The invoice is downstream of that discipline.
Sources
- Stripe Docs, "Usage-based billing": https://docs.stripe.com/billing/subscriptions/usage-based
- Stripe Docs, "Idempotent requests": https://docs.stripe.com/api/idempotent_requests
- AWS, "Analyzing your costs and usage with AWS Cost Explorer": https://docs.aws.amazon.com/cost-management/latest/userguide/ce-what-is.html
- Cloudflare, "GraphQL Analytics API": https://developers.cloudflare.com/analytics/graphql-api/
- OpenMeter documentation: https://openmeter.io/docs/
- OWASP, "Secrets Management Cheat Sheet": https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
Top comments (0)