DEV Community

IversonBlake8417
IversonBlake8417

Posted on

Invoice Speech-to-Text: How One Key Detects 4 Provider Fallback Gaps

TL;DR: Treat an API key as authorization, not evidence that speech-to-text exists. For voice attachments that accompany supplier invoices, expose four decisions for every tenant: capability, region, policy, and selection. Probe each configured adapter with a tiny non-sensitive fixture, cache the result briefly, route only inside the tenant's allowed region, and record usage in neutral units such as audio seconds and bytes. A model list helps discovery. It cannot prove that an audio call will succeed.

Approach Evidence collected Pick this when Main limit
Static configuration Declared region and transcription support Contracts and deployments change rarely Configuration can drift from runtime behavior
Model-list inspection Advertised model identifiers You need a cheap hint before probing A name does not prove access or accepted media
Active capability probe A real result for a known fixture Correct routing matters more than one startup call Probes need caching, timeouts, and safe test data
Request-time fallback Actual result from each eligible adapter Intermittent failures must be absorbed Blind retries can duplicate work and hide cost

The practical default is active probing plus a short-lived cache, with request-time fallback restricted by region and feature flags. Keep static configuration as the policy boundary. Each piece has one job.

Why can one compatible key support text but not speech?

"OpenAI-compatible" often describes a subset of an API surface. Successful authentication, text generation, or model listing does not establish that the same credential, deployment, and base URL accept audio. The failure is easy to misread: the credential is valid, while the capability may be missing, inaccessible, or represented differently.

Model-list detection is still useful. It can eliminate an obviously ineligible candidate without uploading audio. Keep it as a hint because identifiers do not establish media limits, authorization, regional processing, or transcription behavior. The decisive check is a bounded call through the adapter contract production uses.

This distinction matters in invoice intake. A spoken note may contain the purchase-order number, tax treatment, or delivery exception needed by field extraction. Sending that note to an arbitrary fallback can violate a tenant's data-location policy even if the transcript is accurate.

Fast isn't enough.

Pick each strategy when its evidence matches the risk

Use static declarations for hard constraints: an adapter's deployment region, the tenant regions it may serve, and whether operations has enabled it. These are policy, not observations. A runtime probe must never override them.

Use model discovery as a low-cost filter when an adapter exposes it. Normalize the response behind your own interface and look for an explicit audio capability supplied by configuration or the adapter implementation. Don't guess from a model name. The text transcribe can be absent from a valid identifier or present in one that the current credential cannot call.

Use active probes before admitting a candidate to the routing pool. The fixture should be short, synthetic, and free of supplier or customer data. Check more than HTTP success: require non-empty text, record latency, and expire the observation. A cached success from last week says little about today's deployment.

Probe first.

Use request-time fallback only among candidates that already passed policy and capability checks. Retry the same adapter for a transient transport failure only when request semantics prevent duplicate downstream writes. Otherwise, move once to the next eligible candidate and preserve one correlation ID across the attempt chain.

Implement the probe and selector

This TypeScript core has no vendor routes or SDK assumptions. Each adapter owns its wire format; the router owns policy, evidence, and selection.

type Region = "eu" | "us";
type ProbeState = "supported" | "unsupported" | "unavailable";

type Transcript = {
  text: string;
  audioSeconds: number;
  inputBytes: number;
};

interface SpeechAdapter {
  id: string;
  region: Region;
  listCapabilities(): Promise<ReadonlySet<string>>;
  transcribe(audio: Uint8Array, idempotencyKey: string): Promise<Transcript>;
}

type Candidate = {
  adapter: SpeechAdapter;
  enabled: boolean;
  priority: number;
};

type Probe = {
  providerId: string;
  region: Region;
  state: ProbeState;
  checkedAt: number;
  latencyMs: number;
  reason: string;
};

const withTimeout = async <T>(work: Promise<T>, timeoutMs: number): Promise<T> => {
  const timeout = new Promise<never>((_, reject) => {
    setTimeout(() => reject(new Error("probe_timeout")), timeoutMs);
  });
  return Promise.race([work, timeout]);
};

async function probe(
  adapter: SpeechAdapter,
  fixture: Uint8Array,
  now = Date.now(),
): Promise<Probe> {
  const started = Date.now();
  try {
    const capabilities = await withTimeout(adapter.listCapabilities(), 1_500);
    if (!capabilities.has("audio.transcription")) {
      return {
        providerId: adapter.id,
        region: adapter.region,
        state: "unsupported",
        checkedAt: now,
        latencyMs: Date.now() - started,
        reason: "capability_not_advertised",
      };
    }

    const result = await withTimeout(
      adapter.transcribe(fixture, `probe-${now}`),
      4_000,
    );
    const supported = result.text.trim().length > 0;
    return {
      providerId: adapter.id,
      region: adapter.region,
      state: supported ? "supported" : "unsupported",
      checkedAt: now,
      latencyMs: Date.now() - started,
      reason: supported ? "fixture_transcribed" : "empty_transcript",
    };
  } catch (error) {
    return {
      providerId: adapter.id,
      region: adapter.region,
      state: "unavailable",
      checkedAt: now,
      latencyMs: Date.now() - started,
      reason: error instanceof Error ? error.message : "unknown_error",
    };
  }
}

type TenantPolicy = {
  tenantId: string;
  allowedRegions: ReadonlySet<Region>;
  speechEnabled: boolean;
};

function selectCandidates(
  candidates: Candidate[],
  probes: Map<string, Probe>,
  policy: TenantPolicy,
  now = Date.now(),
): Candidate[] {
  const maxProbeAgeMs = 5 * 60 * 1_000;
  if (!policy.speechEnabled) return [];

  return candidates
    .filter(({ adapter, enabled }) => {
      const evidence = probes.get(adapter.id);
      return (
        enabled &&
        policy.allowedRegions.has(adapter.region) &&
        evidence?.state === "supported" &&
        now - evidence.checkedAt <= maxProbeAgeMs
      );
    })
    .sort((a, b) => a.priority - b.priority);
}

type AttemptEvent = {
  tenantId: string;
  invoiceId: string;
  providerId: string;
  region: Region;
  outcome: "success" | "failure";
  audioSeconds: number;
  inputBytes: number;
  latencyMs: number;
  reason?: string;
};

async function transcribeWithFallback(
  eligible: Candidate[],
  policy: TenantPolicy,
  invoiceId: string,
  audio: Uint8Array,
  emit: (event: AttemptEvent) => void,
): Promise<Transcript> {
  for (const { adapter } of eligible) {
    const started = Date.now();
    try {
      const result = await adapter.transcribe(
        audio,
        `${policy.tenantId}:${invoiceId}`,
      );
      emit({
        tenantId: policy.tenantId,
        invoiceId,
        providerId: adapter.id,
        region: adapter.region,
        outcome: "success",
        audioSeconds: result.audioSeconds,
        inputBytes: result.inputBytes,
        latencyMs: Date.now() - started,
      });
      return result;
    } catch (error) {
      emit({
        tenantId: policy.tenantId,
        invoiceId,
        providerId: adapter.id,
        region: adapter.region,
        outcome: "failure",
        audioSeconds: 0,
        inputBytes: audio.byteLength,
        latencyMs: Date.now() - started,
        reason: error instanceof Error ? error.message : "unknown_error",
      });
    }
  }
  throw new Error("no_eligible_transcription_adapter");
}
Enter fullscreen mode Exit fullscreen mode

The diagram in words is short: tenant policy narrows the region; the feature flag opens or closes speech intake; fresh probes narrow the adapters; priority orders the survivors; the attempt ledger explains the final choice. No signal does two jobs.

Notice the three probe outcomes. unsupported means the adapter gave evidence against the capability. unavailable means the test couldn't decide. Collapsing both into false makes an outage look like permanent incompatibility and can keep a recovered adapter out of service.

The 1.5-second discovery timeout, 4-second transcription timeout, and five-minute cache are example operating parameters. They aren't universal recommendations. Set them from your latency objective and probe volume; a copied number can quietly become policy.

Make tenant cost visible without making price the router

A per-tenant ledger should answer which invoice caused work, which adapter performed it, where it ran, how much audio was processed, and whether fallback added an attempt. Store audioSeconds, inputBytes, request count, and latency before applying any billing formula. Those measurements remain useful when contracts change. Don't put mutable unit prices in application code. Join usage to a versioned rate table in the reporting layer, with effective timestamps and currency recorded beside the result. If finance changes an allocation rule, historical raw usage stays intact and the report can be recomputed. Keep invoice identifiers pseudonymous in telemetry where possible. Logs need routing facts, not supplier names or transcript contents. HIPAA isn't automatically applicable to supplier invoices, but its access-control, audit-control, integrity, and transmission-security requirements are useful prompts when the same intake service can process protected health information. Applicability is a legal and data-classification question, not a library setting.

Keep both records.

For live operator views, Server-Sent Events can carry probe-state changes and aggregated counters over a persistent HTTP connection. The browser EventSource interface reconnects by default when a connection closes, and the stream uses text/event-stream. Don't stream transcripts. Send bounded operational events, authorize the stream, and keep the durable ledger behind it.

A useful alert is specific: eligible adapter count reached zero for an enabled tenant region. A weaker alert fires on every failed attempt, including failures the fallback path handled. Page on lost service. Graph the handled failures.

Roll out the feature flag in four checks

Start with the global speech flag off. Enable probe traffic for one region with the synthetic fixture, but don't route invoice audio yet. Once capability observations are fresh, enable a small set of internal tenants whose region policy is explicit. Then expand by tenant cohort while watching zero-eligible counts, fallback depth, transcript emptiness, and usage per invoice.

The flag belongs at the tenant-and-region boundary. A single global boolean cannot express "EU processing allowed, US processing denied" for one tenant and the reverse for another. Record flag evaluation with a configuration version so an operator can reconstruct why a request was admitted.

Rollback is plain: close admission for new speech jobs, let accepted jobs finish under their captured configuration version, and keep the text-field extraction path available for invoices without audio. This prevents a runtime capability issue from stopping the entire invoice workflow.

Limits

A successful probe does not certify transcript quality, language coverage, media limits, or regulatory suitability. Test those separately with representative, permitted data and explicit acceptance thresholds. The selector also assumes adapters implement idempotent or otherwise duplicate-safe behavior; verify that contract before adding retries.

Static policy can be too rigid for tenants that permit several regions, while request-time fallback adds latency and extra processing attempts. Model discovery may be cheaper than probing, but it supplies weaker evidence. These trade-offs are why the decision trail matters: policy said where, the flag said whether, the probe said what worked, and the ledger said what each tenant consumed.

References

Top comments (0)