TL;DR: Pick a gateway for moderation classification by the evidence it preserves, not by how many model names fit behind one key. Require every provider path to produce the same small JSON contract, validate it before accepting the result, and record why routing or fallback happened. A single application credential can simplify integration, but it does not make provider outputs interchangeable.
| Approach | Pick this when | Main operational trade-off | Correctness check |
|---|---|---|---|
| Direct provider adapters | The provider set is small and the team wants explicit control | More credentials and adapter code | Validate after each adapter normalizes its response |
| Managed aggregation gateway | One application integration and centralized policy matter most | Another trust and failure boundary | Validate the gateway result independently in the application |
| Self-hosted routing layer | Routing rules, telemetry, and data handling need local control | The team owns upgrades and on-call work | Validate at ingress and again after every provider response |
This field guide is about a media queue where users report posts, comments, or messages before a human reviews them. The model assigns a narrow category and urgency. It does not remove content. That distinction keeps the architecture honest: the classifier prioritizes work, while deterministic policy and people retain authority over consequential action.
Keep that boundary visible.
What should one API key preserve across multiple LLM providers?
One gateway key gives the application one authentication boundary. It can also give the team one place to express routing policy, attach request identifiers, and meter traffic. Those are useful properties. They are not proof that two providers interpret a schema, refusal, timeout, or safety boundary in the same way.
Think of the path in five boxes: report envelope, policy router, provider adapter, schema validator, human-review queue. The arrow from the adapter to the validator is the critical one. No result crosses into the queue until it is parsed and checked. If a fallback model runs, its output travels through that same arrow. There is no trusted shortcut.
This framing changes the comparison. Provider count is inventory. The engineering question is whether the gateway exposes enough information to distinguish a transport failure from a valid but unusable classification. A 200 response containing an unknown label is still a failed classification. So is syntactically valid JSON with a missing report identifier.
Those failures count.
Keep the application credential in a secret manager and rotate it under the gateway's documented process. Upstream credentials, if the gateway uses them, remain separate assets with separate scopes and audit trails. “One key” describes the caller experience, not the whole security model.
Direct adapters expose the whole provider matrix
Direct adapters are a serious option for two or three provider paths. They make the boundaries visible: each adapter owns authentication, timeout translation, response extraction, and normalization into an internal result. There is less ambiguity during an incident because the application knows which upstream it called.
The cost is repetition. Every adapter must implement cancellation, bounded retries, error taxonomy, and telemetry with the same semantics. Provider-native structured-output features can improve conformance, but the application still validates the received value. JSON Schema defines a vocabulary for describing instance structure; it does not make a remote response trustworthy merely because a request included a schema.
Choose this path when a small team can review every adapter change and provider diversity moves slowly. It is also a clean baseline for evaluating a gateway: run the same fixed report corpus through both paths and compare contract-valid rate, label disagreement, refusal rate, and end-to-end latency distributions. Do not collapse those signals into one “success” counter.
Aggregation creates a central evidence boundary
An aggregation gateway earns its place when centralized policy is the job: tenant quotas, allowed-model lists, routing constraints, credential isolation, or a consistent request envelope. The application should depend on its own interface rather than a gateway-specific response shape. That keeps the moderation workflow readable and makes replacement a bounded adapter task.
The important comparison is semantic visibility. Can operators tell which provider and model handled a report? Is the routing reason available? Are timeout, rate-limit, refusal, malformed JSON, and schema rejection distinguishable? Can a request ID be correlated across the application and gateway without logging user text? If those answers are vague, one key has hidden the evidence needed to operate the system.
Fallback also needs a written rule. For this workload, fallback is appropriate after a retryable transport or availability failure, or after a response fails the classification contract. It should not silently reinterpret an accepted result merely because another model might choose a different label. That would turn resilience into nondeterministic adjudication.
Avoid a cheapest-first router as the default decision rule. Price is only one constraint, and a mutable per-token number says nothing about schema-valid rate on the team's actual moderation taxonomy. A better router filters candidates by data-handling requirements and observed contract performance, then applies latency and budget ceilings. Human-review capacity belongs in that calculation: a low-confidence flood can cost attention even when inference spend is modest.
Cheap invalid output is waste.
Self-hosting turns routing evidence into owned infrastructure
A self-hosted layer gives the team direct control over logs, sampling, rollout, and the normalization contract. It fits environments where report data must stay within a defined network path or where routing decisions need custom, reviewable logic.
Ownership is real. Someone must patch dependencies, protect credentials, cap retries, shed load, and test provider changes. The on-call team now owns the layer that sits between every report and every classifier. Pick it because that control is required, not because the proxy looks small in a diagram.
Start with boring routing: an allowlist, a primary candidate, one eligible fallback, and a total deadline. Add adaptive routing only after the telemetry can explain it. Clever selection without legible evidence is hard to debug and harder to defend during a moderation-policy review.
Deep implementation: validate before fallback
The internal contract should be smaller than any provider's native response. Four fields are enough for this example: the report identifier, one allowed queue label, an urgency integer, and a short reason for the human reviewer. The label set comes from the review operation, not from a model's preferred vocabulary.
Here is a minimal Node.js boundary. It uses generic adapters, checks unknown keys, and separates retryable provider failure from invalid output. All candidates face the same parser.
type QueueLabel = "threat" | "harassment" | "spam" | "other";
type Classification = {
reportId: string;
label: QueueLabel;
urgency: 1 | 2 | 3 | 4 | 5;
reason: string;
};
type Candidate = {
id: string;
classify(input: string, signal: AbortSignal): Promise<unknown>;
};
const labels = new Set<QueueLabel>([
"threat",
"harassment",
"spam",
"other",
]);
function parseClassification(value: unknown, reportId: string): Classification {
if (typeof value !== "object" || value === null || Array.isArray(value)) {
throw new Error("schema:root");
}
const record = value as Record<string, unknown>;
const allowedKeys = new Set(["reportId", "label", "urgency", "reason"]);
if (Object.keys(record).some((key) => !allowedKeys.has(key))) {
throw new Error("schema:unknown_key");
}
if (record.reportId !== reportId) throw new Error("schema:report_id");
if (typeof record.label !== "string" || !labels.has(record.label as QueueLabel)) {
throw new Error("schema:label");
}
if (!Number.isInteger(record.urgency) || Number(record.urgency) < 1 || Number(record.urgency) > 5) {
throw new Error("schema:urgency");
}
if (typeof record.reason !== "string" || record.reason.length < 1 || record.reason.length > 240) {
throw new Error("schema:reason");
}
return record as Classification;
}
Notice what is absent: a free-form category, an inferred identifier, and automatic enforcement. Strict rejection is deliberate. Mapping an unknown label to other would make a broken contract look healthy, while copying the input ID into a response that omitted it would conceal an association failure.
The router adds a total deadline and records one event per attempt. In production, emitAttempt would send a metric and a span event through the application's telemetry layer. It should never include the report body.
type AttemptEvent = {
requestId: string;
candidateId: string;
outcome: "accepted" | "timeout" | "provider_error" | "schema_rejected";
elapsedMs: number;
};
declare function emitAttempt(event: AttemptEvent): void;
export async function classifyWithFallback(
requestId: string,
reportId: string,
reportText: string,
candidates: readonly Candidate[],
totalTimeoutMs = 8_000,
): Promise<Classification> {
const deadline = Date.now() + totalTimeoutMs;
for (const candidate of candidates) {
const remainingMs = deadline - Date.now();
if (remainingMs <= 0) throw new Error("classification_deadline_exceeded");
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), remainingMs);
const startedAt = performance.now();
try {
const raw = await candidate.classify(reportText, controller.signal);
const result = parseClassification(raw, reportId);
emitAttempt({
requestId,
candidateId: candidate.id,
outcome: "accepted",
elapsedMs: performance.now() - startedAt,
});
return result;
} catch (error) {
const message = error instanceof Error ? error.message : "provider_error";
const outcome = controller.signal.aborted
? "timeout"
: message.startsWith("schema:")
? "schema_rejected"
: "provider_error";
emitAttempt({
requestId,
candidateId: candidate.id,
outcome,
elapsedMs: performance.now() - startedAt,
});
} finally {
clearTimeout(timer);
}
}
throw new Error("classification_unavailable");
}
This sample intentionally gives all attempts a shared eight-second budget. A tempting first design gives every fallback a fresh timeout because each adapter looks independent. Follow that path with three candidates: the primary waits eight seconds, the first fallback receives another eight, and the second fallback receives eight more. The report can now wait roughly 24 seconds before network overhead, queue time, and serialization are counted. Meanwhile, a caller may retry because its own deadline expired, creating duplicate work with a different request ID. A single absolute deadline prevents that multiplication. Each new attempt receives only the remaining time, and the queue gets one clear terminal outcome when the budget is exhausted. The trade-off is blunt but explainable: a slow primary leaves less time for its fallback. If that is unacceptable, change the candidate order or the total service objective; do not hide another full timeout inside the adapter.
One clock. One outcome.
The production queue should store the accepted classification with requestId, candidate identity, model identifier if available, contract version, and policy version. Store the original report under the media system's access controls rather than duplicating it into telemetry. OpenTelemetry's trace model supplies trace and span identifiers for correlation, while its guidance warns against placing sensitive data in attributes. Low-cardinality counters can then track attempts by candidate and outcome. Latency belongs in histograms. Alerts should combine sustained schema-rejection rate with queue age; either signal alone can be noisy.
Test this boundary in layers. Unit tests feed the parser missing keys, extra keys, wrong identifiers, fractional urgency values, and oversized reasons. Adapter contract tests replay recorded, redacted response shapes. An offline evaluation set measures label disagreement and confusion by policy category, with human-reviewed expected outcomes. Finally, a staged deployment shadows a small, policy-approved sample without changing queue priority. The release gate is contract validity plus review quality, not JSON parse rate.
Limits that stay visible
Schema-valid output can still be wrong. It can reflect ambiguous taxonomy, weak instructions, adversarial text, or a distribution shift in reports. Confidence text generated by the same model is not independent evidence. Use human-reviewed evaluation data and queue outcomes to judge quality.
Fallback can improve availability while changing behavior. Record it. Keep the candidate list short, bound the total deadline, and surface exhaustion as an unavailable classification that enters a safe manual queue. Never turn failure into an invented label.
A gateway also cannot erase data-governance obligations. Teams still need to review retention, regional processing, subprocessor boundaries, incident procedures, and which report fields may leave the application. The final choice is therefore conditional: use direct adapters for a small explicit matrix, aggregation for centralized policy, or self-hosting for local control. In every case, the durable design is the same internal contract and the same observable validation boundary.
References
- https://json-schema.org/draft/2020-12/json-schema-core
- https://opentelemetry.io/docs/specs/otel/trace/api/
- https://opentelemetry.io/docs/security/handling-sensitive-data/
- https://owasp.org/www-project-application-security-verification-standard/
- https://platform.openai.com/docs/guides/function-calling
- https://openrouter.ai/docs
Top comments (0)