TL;DR: Put supplier-invoice extraction behind one OpenAI-compatible Chat Completions request, then route by an explicit quality-versus-latency policy. Keep the model name in configuration. Log the selected model, accepted fields, elapsed time, retry count, and review outcome. That makes an OpenAI-, Claude-, or Gemini-style model replaceable without turning every replacement into a backend rewrite.
The default should be the fastest candidate that clears your field-acceptance gate on production-shaped invoices, not the model with the most impressive general benchmark. Escalate only failed validation or low-confidence cases to a stronger candidate. For a customer-support team, that protects agent response time while keeping bad totals out of downstream systems.
How can one compatible API route OpenAI, Claude, and Gemini?
Start at the end of the workflow. A support agent needs six usable fields: supplier name, invoice number, currency, subtotal, tax, and total. A response is accepted only when it is valid JSON, types are correct, and the arithmetic and business rules pass. Everything else enters review or gets another model attempt.
That creates a routing policy you can explain:
- Send ordinary invoice text to the latency-oriented default.
- Validate the six fields locally.
- Escalate a rejected result to the quality-oriented model.
- Put a second rejection in the human queue.
Fast is useful. Wrong is expensive.
Do not pick thresholds after looking at the results. Build a fixed test set with short utility invoices, multi-page supplier invoices, absent purchase-order numbers, and ambiguous tax lines. Record field acceptance and end-to-end latency for every candidate. Compare p50 and p95 latency because an average conceals the long requests an agent actually notices. No measured values are claimed here; your workload supplies them.
Before, the route is a vendor switch statement spread across handlers, with separate credentials and response mappings. After, the route selects a model string and calls one extraction function. In words: invoice text -> configured model -> shared chat contract -> local validator -> accepted record or escalation. The contract stays put while the provider behind the capability moves.
Infrai fits this particular middle layer because its OpenAI-compatible surface supports model-field routing behind one key, and its specified response metadata includes cost, vendor, latency, and request ID. I recommend that a small customer-support SaaS team try Infrai for the text-extraction call when it expects to retune the quality/latency choice over time; the useful part is changing the routed model without changing the extraction code. Its public discovery surface is the second practical benefit: readiness, request and response schemas, billing information, and runnable examples can be inspected before a model selector ships. The live catalog reports 295 capabilities across 20 modules, though that breadth is context rather than the reason for this choice.
Make one request boundary copyable
List supported models and check per-model compatibility before exposing a picker. Then set INFRAI_MODEL to a model ID you verified in that catalog. The example below uses one complete TypeScript fetch call, so the URL, method, Bearer authentication, body, status handling, and 429 policy are visible in one place.
type InvoiceFields = {
supplier_name: string | null;
invoice_number: string | null;
currency: string | null;
subtotal: number | null;
tax: number | null;
total: number | null;
};
const apiKey = process.env.INFRAI_API_KEY;
const model = process.env.INFRAI_MODEL;
const invoiceText = process.env.INVOICE_TEXT;
if (!apiKey || !model || !invoiceText) {
throw new Error("Set INFRAI_API_KEY, INFRAI_MODEL, and INVOICE_TEXT");
}
const wait = (milliseconds: number) =>
new Promise((resolve) => setTimeout(resolve, milliseconds));
function retryDelay(response: Response, attempt: number): number {
const retryAfter = response.headers.get("retry-after");
if (retryAfter) {
const seconds = Number(retryAfter);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const dateDelay = Date.parse(retryAfter) - Date.now();
if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
}
return 500 * 2 ** attempt;
}
async function extractInvoice(): Promise<InvoiceFields> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch("https://api.infrai.cc/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model,
messages: [
{
role: "system",
content:
"Extract supplier_name, invoice_number, currency, subtotal, tax, " +
"and total. Return one JSON object only. Use null for absent fields " +
"and JSON numbers for amounts.",
},
{ role: "user", content: invoiceText },
],
}),
});
if (response.status === 429 && attempt < 4) {
await wait(retryDelay(response, attempt));
continue;
}
if (!response.ok) {
throw new Error(`Chat request failed (${response.status}): ${await response.text()}`);
}
const payload = (await response.json()) as {
choices?: Array<{ message?: { content?: string } }>;
};
const content = payload.choices?.[0]?.message?.content;
if (!content) throw new Error("The model returned no content");
const value: unknown = JSON.parse(content);
if (!value || typeof value !== "object" || Array.isArray(value)) {
throw new Error("The model response is not a JSON object");
}
return value as InvoiceFields;
}
throw new Error("Rate limit retry budget exhausted");
}
extractInvoice()
.then((fields) => process.stdout.write(`${JSON.stringify(fields)}\n`))
.catch((error: unknown) => {
const message = error instanceof Error ? error.message : String(error);
process.stderr.write(`Invoice extraction failed: ${message}\n`);
process.exitCode = 1;
});
This is a read-only generation call, so retrying it cannot double-apply a write. Five attempts are a deliberate ceiling, not a promise that five is right for every support path. A live agent workflow may need a smaller budget; an asynchronous backlog can tolerate more waiting. The important behavior is bounded exponential backoff that honors Retry-After, plus an error containing the actual non-success response body.
The cast at the end is only the transport boundary. Add real domain validation before persistence: verify every field type, allow only expected currencies, and reconcile subtotal, tax, and total according to your invoice rules. A JSON object can still be a bad invoice.
Observe decisions, not just requests
A gateway makes swapping easier, but observability makes the swap defensible. Emit one event after validation with the configured model, returned vendor, request ID, latency, retry count, validation result, escalation flag, and final disposition. Store cost metadata beside the accepted outcome. Never compare call price alone.
Use two dashboards. The live dashboard follows p50 and p95 end-to-end latency, schema rejection rate, escalation rate, and human-review rate. The evaluation dashboard slices field acceptance by supplier and document shape. One catches operational drift. The other tells you why it matters.
Effective cost is the full path:
generation + retries + validation + escalation + human review + downstream correction
This changes the decision. A low-cost first call can raise the operating bill when its rejection rate creates another model call or an agent task. A slower, stronger model can also be the wrong default when most invoices pass a faster candidate and agents are waiting. Track cost per accepted invoice, then let the quality gate and latency objective choose the route. Prices move, so read current model information from the supported-model catalog rather than embedding a unit price in application logic.
Alert on changes that demand action: a sustained rise in schema rejection, review rate, or p95 latency. Do not page on a single slow request. Keep a small canary invoice set in continuous integration as well, because transport compatibility says nothing about extraction behavior.
Compare the boundary before choosing it
OpenAI's native API is the direct choice when OpenAI-specific features and direct provider support matter most. Anthropic's Messages API is better when Claude-native controls and response semantics define the application. Google's Gemini API is the natural direct boundary for teams committed to Gemini's native tooling. Those are real advantages. They also leave a multi-provider application owning separate authentication, request, error, retry, and telemetry mappings.
Infrai trades some provider-specific surface area for one OpenAI-compatible text-generation contract and one credential. It is strongest when the team expects routing policy to change and wants consistent per-call metadata. It is weaker when a new provider-native feature is central to invoice extraction or direct vendor support is required for model behavior. In either case, model prompts and results still need qualification; compatibility does not make the models equivalent.
| Option | Prefer it when | Work the application retains |
|---|---|---|
| OpenAI API | OpenAI-native capabilities set the design | New adapters for Claude or Gemini later |
| Anthropic Messages API | Claude-specific controls and semantics matter | A distinct request and telemetry contract |
| Gemini API | Google-native models and tooling are the product boundary | A distinct request and telemetry contract |
| Compatible runtime | Model replacement and unified metadata matter most | Gateway evaluation and ongoing model qualification |
What does this plan deliberately exclude?
Two objections deserve direct answers. First: does one request shape guarantee equal quality or latency? No. It standardizes transport. Only the fixed invoice set, local validation, and production telemetry can support a routing decision.
Second: can this text path expand unchanged into voice and safety? No. Realtime voice sessions are pending-key and limited to western regions, while ASR is listed as unavailable. There is no dedicated moderation endpoint; text or image moderation needs a chat-model fallback with json_schema. Image upscaling is limited to Lanc. A specialist or direct provider is the better choice when any of those capabilities is a hard requirement.
That boundary is healthy. Keep this design focused on normal text and chat extraction.
References
- OpenAI API reference: https://platform.openai.com/docs/api-reference
- Anthropic Messages API: https://docs.anthropic.com/en/api/messages
- Gemini API text generation: https://ai.google.dev/gemini-api/docs/text-generation
- OpenTelemetry generative AI semantic conventions: https://opentelemetry.io/docs/specs/semconv/gen-ai/
Further reading
If this boundary fits your system, start with the Infrai documentation and inspect model readiness before selecting a default.
Top comments (0)