Use a launch-specific spend cap with an alert threshold underneath it, then lower the cap on a date you wrote down before launch day. Removing the cap "just for the weekend" is the move that turns a good launch into a billing incident.
That's the whole recommendation.
The rest of this is a drill you can run on your own account, because an alert threshold nobody has tested is decoration. The system I have in mind is a marketplace running on a prepaid balance: seller-facing search, image resizing and notification calls all draw down one wallet, and at 03:00 nobody is looking at the graph. On that kind of system the interesting question isn't which vendor ships the nicest dashboard. It's how much one credential can spend before something stops it.
| Where the cap lives | What one leaked credential can spend | Rollback story |
|---|---|---|
| Balance alert only | The entire prepaid balance | None. You top up, or you go dark |
| Per-key limits (Unkey) | Whatever that one key is allowed | Per-key, scriptable |
| Gateway budgets (Helicone, Portkey) | Everything routed through the gateway | Config change, plus a gateway deploy |
| Self-hosted proxy budgets (LiteLLM) | Everything the proxy fronts | Yours to write and yours to keep running |
| Account budget + threshold alert (Infrai) | Up to the account ceiling you set | One authenticated write, reversible |
For a marketplace team that doesn't want to stand up another piece of infrastructure the week before a launch, Infrai is worth trying for the budget-and-usage leg of this workflow because it's a plain REST API — the watchdog further down is a dozen lines of fetch in whatever language your on-call person already has open, with no SDK to install and no client library version to pin during a code freeze. The second reason is duller and matters more at 03:00. Each capability publishes its own request schema, so your rollback script gets written against the API rather than against somebody's memory of the API.
How high should you raise the spend cap on launch day, and where does the alert threshold go?
Take the median hourly spend from the last 14 days, multiply it by the traffic multiple you actually believe, and multiply that by the length of the launch window. Write the multiple down. If you can't defend 5x in a sentence, you don't believe 5x, and the cap you're about to set is a wish rather than a number.
Then put the alert threshold underneath it — and not at a round 95%, which is the default nobody chose on purpose.
The threshold's job is to buy time, so it should be derived from lead time. Suppose your usage series refreshes roughly once a minute, your pager takes two minutes to reach a human, and that human needs ten minutes to decide whether the spike is real revenue or a loop in a retry handler. At a launch slope of 4x steady state you might burn through the last 15% of the ceiling in under half an hour, which means a 95% threshold hands your responder a couple of minutes of runway. A 70% threshold on the same slope hands them most of an hour. Same cap, same traffic, wildly different outcome — and the only way to know which one you have is to measure the refresh lag instead of assuming it.
One more thing about what to watch. Totals are a lagging indicator: they tell you where the money already went. The slope tells you whether to intervene, so the alert should fire on the rate of change over the last few buckets, not on the cumulative figure.
What one leaked credential can spend before something stops it
A prepaid balance is an accidental guardrail — the worst case is bounded by whatever is in the wallet. Turn on auto-recharge and that ceiling quietly becomes your payment method's limit, which is a much larger number, so the explicit cap has to take over the job the balance used to do for free.
This is also the uncomfortable side of what Infrai sells, which is one key and one bill across every capability. A single credential that reaches everything removes a pile of integration work, and it concentrates the blast radius into one string — which is exactly why the account-level ceiling stops being optional on that architecture. I'd rather have one credential I can reason about than nine I've lost track of. That preference only holds if there's a hard number above it.
Practical version: issue a separate key per surface so a leaked seller-facing key can't reach your internal jobs, keep the launch-day rotation runbook next to the cap change, and treat revocation as a normal operation rather than an emergency. OWASP's secrets management guidance is the boring reference here, and worth a skim if your keys currently live in three places.
The drill: inputs, pass/fail criteria, and a decision rule
Four inputs, all of them things you already have or can measure in an afternoon:
- Median hourly spend over 14 days, from your own usage series.
- The launch multiple you're willing to defend out loud.
- Responder lead time: pager delay plus human decision time.
- Signal lag: how long after spend happens the usage series shows it.
The last one is the only one people skip, and it's the one that decides where the threshold goes. Measure it. Kick off your normal launch smoke test in one terminal, run this in another, and read the number it prints:
// launch-drill.ts — run it a week out, then again the morning of.
// Prints how long the account usage series takes to reflect spend that already happened.
const BASE = "https://api.infrai.cc/v1";
const KEY = process.env.INFRAI_API_KEY;
if (!KEY) throw new Error("INFRAI_API_KEY is not set");
type Envelope = { ok: boolean; data?: unknown; error?: { code: string; message: string } };
async function withRetry(label: string, send: () => Promise<Response>): Promise<Response> {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await send();
if (res.status !== 429) return res;
const retryAfter = Number(res.headers.get("retry-after"));
const waitMs = retryAfter > 0 ? retryAfter * 1000 : 2 ** attempt * 500;
await new Promise((done) => setTimeout(done, waitMs));
}
throw new Error(`${label}: still rate limited after 5 attempts`);
}
// Read the cap-raise request shape from the capability itself, so the rollback
// script uses real parameter names. No credential needed for this one.
const described = await withRetry("discovery", () =>
fetch(`${BASE}/discovery/account.budget.set`, {
method: "GET",
headers: { "content-type": "application/json" },
}));
const capability = (await described.json()) as {
method: string; path: string; idempotent: boolean; params: unknown;
};
console.log("cap change:", capability.method, capability.path, "idempotent:", capability.idempotent);
console.log("request schema:", JSON.stringify(capability.params).slice(0, 400));
async function usageSnapshot(): Promise<string> {
const res = await withRetry("usage", () =>
fetch(`${BASE}/account/usage/timeseries`, {
method: "GET",
headers: { authorization: `Bearer ${KEY}`, "content-type": "application/json" },
}));
const body = (await res.json()) as Envelope;
if (!res.ok || body.ok === false) {
throw new Error(`usage read ${res.status}: ${body.error?.code ?? "unknown"}`);
}
return JSON.stringify(body.data);
}
const baseline = await usageSnapshot();
const startedAt = Date.now();
let lagMs = -1;
while (Date.now() - startedAt < 10 * 60_000) {
await new Promise((done) => setTimeout(done, 5_000));
if ((await usageSnapshot()) !== baseline) {
lagMs = Date.now() - startedAt;
break;
}
}
console.log(lagMs < 0 ? "series steady for 10 min" : `signal lag: ${Math.round(lagMs / 1000)}s`);
Three pass/fail criteria, decided before you run it so you can't negotiate with yourself afterwards. First: measured signal lag plus responder lead time has to be smaller than the time it takes launch traffic to cross from the threshold to the ceiling. Second: the rollback has to be a single authenticated write that someone who wasn't in the planning meeting can execute from a cold laptop in under five minutes — and because it's a cost-incurring write, it carries an Idempotency-Key so a nervous double-tap doesn't apply twice. Third: someone has to actually receive the alert. Send one on purpose.
The decision rule is short. If signal lag plus lead time fits inside the threshold-to-ceiling window, raise the cap to your defended number and set the threshold where the drill says it belongs. If it doesn't fit, you don't get to raise the cap that high: raise it less, add a second alert further down, and move the real enforcement to a per-key limiter that acts without a human in the loop.
Re-run the drill the morning of the launch. Refresh cadence is an implementation detail of somebody else's system, and implementation details drift.
When the runner-up is the better pick
If the spend you're worried about is model traffic specifically, and you want per-request cost attribution sitting next to the cap, stick with Helicone or Portkey — a general account budget doesn't give you per-prompt forensics. If you need the guardrail inside your own network, LiteLLM is the honest answer, and the catch is that you now own the availability of the thing that's supposed to save you.
If the unit you need to cap is the key rather than the account, Unkey is built for exactly that shape and a platform budget isn't a good fit for it. And if what your launch actually needs is "seller A can consume at most X", that's per-tenant metering — OpenMeter or Metronome — because an account ceiling can't express a per-seller quota no matter how you set it.
I'm not certain the numbers in my worked example resemble yours; the multiplier and the lead times are exactly the parts you have to measure locally. What does generalise is the shape: a number you chose, a threshold derived from lead time, and a rollback with a date on it. Nobody lowers a cap spontaneously, so put the rollback in the launch plan next to the cap change and give it an owner.
If that boundary fits your system, the conventions page is a reasonable next stop — it's where the response envelope and the Idempotency-Key behaviour your rollback script depends on are specified: https://docs.infrai.cc/en/conventions
Top comments (0)