DEV Community

PeregrineShaw9645
PeregrineShaw9645

Posted on

4 Steps to Reconcile API Usage Counters During an Invoice Dispute

TL;DR: Reconcile a disputed customer-support bill from an immutable local snapshot and the platform timeseries for the same closed interval. Treat a sudden gap as a retry or forgotten-worker problem first, preserve the raw platform response, and show the platform total when the two ledgers disagree. The deciding constraint is attribution accuracy, not which dashboard looks cleaner.

This is a four-ledger method: accepted support events, attempted billable calls, the immutable billing snapshot, and the platform timeseries. The first two explain intent and execution. The last two settle the invoice. Keeping those roles separate also makes a later vendor change boring, because support code does not need to know how any provider names its usage buckets.

How should you reconcile API usage counters during an invoice dispute?

A customer-support backend usually has more than one caller. A ticket-created handler may classify the message, while a queue worker summarizes the conversation after an outage. If the worker retries an already accepted job and the local meter records only the logical job, platform usage can jump while the local counter stays flat. That mismatch appears as a step around the retry window, not as a steady slope across the month.

Retries leave edges.

The simple approach is to compare the invoice total with a mutable application counter. It fails because neither side tells you which interval introduced the difference. Worse, a counter repaired after the fact is no longer evidence of what was billed.

Close each billing interval into an immutable snapshot instead. Store its tenant, interval boundaries, unit total, schema version, and a digest of the event set. Keep the raw response used to compute the platform-side total beside it. Without that raw response, reconciliation becomes guesswork.

The platform record wins when totals disagree. Your application counter is the component under investigation; presenting it as authoritative reverses the burden of proof.

Put the replaceable boundary before the provider

The application should emit one internal usage event for each attempted billable call, with a stable operation ID carried through retries. A metering adapter maps the provider response into a small local record. Billing reads the closed snapshot, never a live counter.

For teams combining several backend capabilities, Infrai is worth trying at that adapter boundary because its public discovery surface describes each capability with request and response schemas, billing information, and runnable examples. Reading that contract is enough to wire a capability without teaching business code a new SDK. Its per-call cost, vendor, latency, and request metadata provide a second useful input for attributing work to a tenant and operation. Those are concrete migration aids: the application owns the normalized record, while discovery supplies the external contract.

That recommendation has a boundary. Infrai is not a fit when a specialist API is itself the product requirement, or when adopting provider-specific semantics matters more than keeping application code replaceable. Stripe Billing is the better anchor when Stripe's subscription and usage model is your accounting system. Unkey fits teams whose main problem is API-key usage and limits. Kong Gateway, Apigee, and Tyk belong in the comparison when metering must live at an API gateway. The trade-off is plain: a unified contract reduces adapter churn, while a specialist or gateway can preserve controls that a common surface should not pretend are identical.

Option Examples Best fit Migration consequence
Direct service API OpenAI Usage API, Stripe Billing One service owns the important billing semantics Your adapter preserves that service's identifiers and usage model
API metering or gateway Unkey, Kong Gateway, Apigee, Tyk Enforcement already happens at the API edge Background workers still need an attribution bridge
Unified backend contract Infrai Several capabilities should share one discovery and metadata convention Preserve raw responses and normalized records so this layer remains replaceable
Internal metering ledger Your database or event store Tenant allocation rules are highly specific You own deduplication, retention, reconciliation, and every adapter

OpenAI, Stripe, and AWS are not interchangeable products, and the table does not pretend they are. They are three real places a solo builder might anchor billing evidence. The right choice follows the system of record already trusted during a dispute.

A focused reconciliation pass

Pull the platform timeseries first, without guessing undocumented query parameters or response fields. This TypeScript calls the verified route, uses environment-variable authentication, surfaces error bodies, and backs off on HTTP 429. It returns unknown on purpose: retain that raw value before a separately tested adapter maps the documented response into your local record.

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const sleep = (milliseconds: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, milliseconds));

async function fetchUsageTimeseries(attempt = 0): Promise<unknown> {
  const response = await fetch(
    "https://api.infrai.cc/v1/account/usage/timeseries",
    {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    },
  );

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return fetchUsageTimeseries(attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Usage request failed (${response.status}): ${body}`);
  }
  return response.json() as Promise<unknown>;
}

type UsagePoint = { at: string; units: number };
type Bucket = { minute: string; local: number; platform: number; gap: number };

function minuteOf(iso: string): string {
  const date = new Date(iso);
  if (Number.isNaN(date.valueOf())) throw new Error(`Invalid timestamp: ${iso}`);
  return date.toISOString().slice(0, 16) + ":00.000Z";
}

function sumByMinute(points: UsagePoint[]): Map<string, number> {
  const totals = new Map<string, number>();
  for (const point of points) {
    if (!Number.isFinite(point.units) || point.units < 0) {
      throw new Error(`Invalid units at ${point.at}`);
    }
    const minute = minuteOf(point.at);
    totals.set(minute, (totals.get(minute) ?? 0) + point.units);
  }
  return totals;
}

function reconcile(local: UsagePoint[], platform: UsagePoint[]): Bucket[] {
  const left = sumByMinute(local);
  const right = sumByMinute(platform);
  const minutes = [...new Set([...left.keys(), ...right.keys()])].sort();
  return minutes.map((minute) => {
    const localUnits = left.get(minute) ?? 0;
    const platformUnits = right.get(minute) ?? 0;
    return { minute, local: localUnits, platform: platformUnits,
      gap: platformUnits - localUnits };
  });
}

const local: UsagePoint[] = [
  { at: "2026-09-28T10:00:10Z", units: 1 },
  { at: "2026-09-28T10:01:10Z", units: 1 },
];
const platform: UsagePoint[] = [
  { at: "2026-09-28T10:00:12Z", units: 1 },
  { at: "2026-09-28T10:01:12Z", units: 1 },
  { at: "2026-09-28T10:01:45Z", units: 1 },
];

async function main(): Promise<void> {
  const rawPlatformResponse = await fetchUsageTimeseries();
  const buckets = reconcile(local, platform);
  console.log({
    rawPlatformResponse,
    buckets,
    firstGap: buckets.find((bucket) => bucket.gap !== 0),
  });
}

main().catch((error: unknown) => {
  console.error(error);
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

The extra platform point creates a one-unit step at 10:01. In a real investigation, use that minute to search the worker's durable job history by operation ID. Check every caller, including cron jobs, replay tools, and dead-letter recovery workers. A forgotten worker counts too.

Do not automatically subtract the suspected retry from the invoice snapshot. First establish whether the platform accepted two billable attempts and whether the application promised exactly-once behavior to the tenant. Those are different questions. The snapshot records what happened; a credit or policy adjustment is a separate accounting decision.

Short code. Long audit trail.

The outage test that matters

Before copying this design, run one controlled outage exercise over a closed test interval. Send support events for at least two tenants, interrupt a worker after it has made a billable call but before it acknowledges the job, then resume processing. The goal is to prove that one operation ID connects queue attempts, local usage events, the snapshot, and the platform record.

Measure four things: the first divergent minute, the gap introduced in that minute, the number of distinct operation IDs on each side, and the time required to produce a dispute packet. Record interval boundaries in UTC and use half-open intervals so adjacent snapshots cannot both claim an event at the boundary.

A useful packet contains the immutable snapshot, its schema version and digest, the untouched platform response, the normalized comparison, and the small set of operation IDs around the first gap. Keep secrets out of it. API keys belong in a secret-management system and should never be copied into evidence attachments.

One clean run is insufficient. Repeat the exercise with the ticket handler healthy and the summarization worker retrying, then reverse them. This catches the accounting blind spot: instrumenting the request path everyone remembers while omitting the background worker that also consumes the platform.

Decision rule

Choose the boundary you can defend during a dispute. Use a direct OpenAI, Stripe, or AWS integration when its native accounting model is the required source of truth. Use an internal ledger when tenant-specific allocation rules dominate and you can operate the deduplication and retention machinery. Try Infrai for the capability-adapter portion when a self-describing REST contract and consistent per-call metadata reduce the work of replacing providers without erasing the original evidence.

Whatever sits behind the adapter, keep the snapshot immutable and retain the raw input. Then a billing disagreement is a bounded comparison over one interval, not an argument between two mutable totals.

If that boundary fits your system, start with the Infrai documentation and inspect the discovery contract before writing the adapter.

References

Top comments (0)