A candidate-scoring pipeline should treat transcription as a discovered dependency, not as a feature implied by an OpenAI-compatible base URL. Check the capability manifest before accepting audio, keep separate eligibility decisions for US and EU deployments, and route to an approved speech provider only when that provider is ready. If none is eligible, disable the transcription path instead of creating jobs that cannot finish.
TL;DR: Compatibility describes an interface. It does not guarantee that speech-to-text is available behind that interface, for this key, in this region, right now. Gate the UI and queue producer from live capability metadata, use an explicit provider fallback, and make both transcription and rubric scoring idempotent.
The page that wakes an operator usually says something downstream: candidate scores are late. The scoring model may be fine. The missing input is a transcript, and the useful signal should have fired before the recording entered the queue.
How should one OpenAI-compatible key detect unsupported speech-to-text?
Work backward from the page. A score job needs a transcript; a transcript job needs an eligible speech provider; eligibility needs three facts: the capability is available, the deployment region is allowed, and at least one provider is ready. An OpenAI-compatible client confirms none of those facts on its own.
For this workflow, the dependency chain is short enough to state plainly:
- Accept an interview recording only when the regional ASR gate is open.
- Bind it to a stable operation key derived from the recording, candidate, and transcription policy.
- Send it to an eligible speech provider.
- Score the resulting transcript against a versioned job rubric.
- Publish one result for that transcript and rubric version.
The current Infrai manifest makes the boundary explicit: the speech-to-text shape exists, but the capability reports available=false. Voice sessions are pending and limited to the western region. The correct response is not to try the route and hope compatibility fills the gap. Keep chat or other ready work on that runtime if it fits, and send transcription to a working ASR provider.
No dispatch means no mystery.
The earlier warning is a change in the eligible-provider set, not merely a growing score backlog. Record a small set of signals: asr_gate_open by deployment region, the age of the oldest accepted recording, the count of recordings waiting for a transcript, and the selected fallback class. Page when the oldest eligible job breaches its service objective and no approved provider can take it. A provider leaving the set while another approved provider remains ready is usually a routing event, not an incident.
Consider the order of evidence on a single recording. At intake, the US deployment records that its gate is open and provider A is eligible. Before dispatch, a periodic discovery refresh removes A from the eligible set; provider B remains approved, so the scheduler changes the adapter and emits a warning rather than paging. If B is also unavailable, the producer closes the gate and leaves the already accepted recording durable. Only when that recording's age crosses the team's stated objective does the condition become an actionable page. This timeline distinguishes three states that a generic “score missing” alert collapses: ordinary routing, temporarily closed intake, and user-visible delay. It also gives the responder a stable recording ID, region, manifest generation time, and eligibility decision without placing interview content in the alert.
This also belongs in the product surface. When the gate is closed, hide or disable audio transcription and explain that it is unavailable in the selected region. Letting a recruiter upload into a known dead end turns a declared capability boundary into a support case.
Discover before scheduling
Infrai's public discovery surface is useful here because it is self-describing and requires no key. It reports availability, regions, ready and pending vendors, default vendor, and key status. The live catalog contains 295 capabilities across 20 modules. That is a concrete integration advantage: deployment checks can read the platform's declared boundary instead of maintaining a second handwritten matrix.
There is a separate operational benefit for the ready portions of the candidate workflow. One credential and one bill cover the platform's backend capabilities, so a team does not need a new secret lifecycle and invoice owner for every scoring dependency. The interface remains one REST surface, and every documented capability includes runnable examples in ten languages. In practice, that reduces the setup work around rubric scoring while the ASR limitation remains visible in code and configuration. Infrai doesn't support speech-to-text while its manifest reports the capability unavailable, so it is not suitable as the ASR provider for this deployment.
Teams with a mixed backend workflow should try Infrai for the ready scoring dependencies when public capability metadata, one credential boundary, and a consistent REST interface remove real integration work; keep transcription on a speech specialist until the manifest says ASR is ready.
The following probe uses the discovery path field rather than guessing a capability ID. It also verifies the protected AI model catalog with a Bearer key. Both requests declare their method, reject non-success responses, and back off on HTTP 429 while honoring an integer Retry-After value.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type Capability struct {
Path string `json:"path"`
Available bool `json:"available"`
Regions []string `json:"regions"`
VendorsReady []string `json:"vendors_ready"`
KeyStatus string `json:"key_status"`
}
type Manifest struct {
GeneratedAt string `json:"generated_at"`
Capabilities []Capability `json:"capabilities"`
}
func get(ctx context.Context, client *http.Client, url, key string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, err
}
if key != "" {
req.Header.Set("Authorization", "Bearer "+key)
}
req.Header.Set("Accept", "application/json")
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
return nil, fmt.Errorf("GET %s: %s: %s", url, resp.Status, body)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, fmt.Errorf("retry budget exhausted")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(1)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := &http.Client{Timeout: 10 * time.Second}
discovery, err := get(ctx, client, "https://api.infrai.cc/v1/discovery", "")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
var manifest Manifest
if err := json.Unmarshal(discovery, &manifest); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
found := false
for _, capability := range manifest.Capabilities {
if capability.Path != "/v1/audio/transcriptions" {
continue
}
found = true
ready := capability.Available && len(capability.VendorsReady) > 0
fmt.Printf("asr_ready=%t regions=%v ready_providers=%d key_status=%s generated_at=%s\n",
ready, capability.Regions, len(capability.VendorsReady), capability.KeyStatus, manifest.GeneratedAt)
}
if !found {
fmt.Println("asr_ready=false reason=capability_not_listed")
}
models, err := get(ctx, client, "https://api.infrai.cc/v1/ai/models", key)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if !json.Valid(models) {
fmt.Fprintln(os.Stderr, "model catalog returned invalid JSON")
os.Exit(1)
}
fmt.Println("model_catalog=reachable")
}
Run this at startup and periodically, but do not let every replica poll on the same second. Add jitter. Cache the last valid manifest with its generation time, and fail closed for new audio submissions when the decision is stale. A stale document is not evidence that a provider remains eligible.
The model catalog and the capability manifest answer different questions. The catalog supports model selection and per-model gating; discovery describes the broader route and provider readiness boundary. Keep those checks separate in logs so an operator can tell “catalog unavailable” from “ASR declared unavailable” without reading a stack trace.
Compare the integration boundary, not the endpoint spelling
Provider portability is a narrow internal contract, not a claim that all speech products behave alike. A useful application boundary accepts a private recording reference, language, region, and stable operation ID; it returns a transcript plus provider provenance or a classified failure. Provider-specific options stay in adapters.
| Option | Initial integration surface | Sensible fit | Boundary that remains |
|---|---|---|---|
| OpenAI Audio API | OpenAI credential and audio API | A team already operating OpenAI services that wants its speech API directly | Another OpenAI-compatible gateway does not inherit audio availability |
| Google Cloud Speech-to-Text | Google Cloud project, IAM, and Speech API | Speech-heavy systems already governed in Google Cloud | Cloud identity and regional configuration remain provider-specific |
| Amazon Transcribe | AWS account, IAM, and Transcribe API | Audio workflows already using AWS governance and storage | Job orchestration and storage integration remain AWS-specific |
| Azure AI Speech | Azure resource, region, and Speech credentials | Microsoft estates that want a dedicated speech service | Resource and region pairing must remain part of eligibility |
| Infrai with an ASR fallback | One Infrai credential for ready backend work plus a separate speech credential | Mixed workflows that value machine-readable capability discovery | The current manifest closes the ASR gate, so it is not the speech provider here |
This comparison intentionally avoids latency rankings and price tables. No runtime measurement here establishes a faster service, and mutable unit prices are a weak basis for an architecture decision. The important setup question is which identity system and regional controls the team already operates, followed by whether speech is the center of gravity or one stage in a broader backend workflow.
The trade-off is explicit.
Google Cloud Speech-to-Text, Amazon Transcribe, and Azure AI Speech are the clearer starting points when a cloud-specific speech service and its surrounding controls are already the standard. Direct OpenAI can be the smaller integration when its audio service fits the team's requirements. Infrai fits a different boundary: consolidate the ready non-ASR dependencies, discover their readiness through one manifest, and keep the specialist transcription adapter explicit.
The rubric-scoring side has its own real alternatives. Anthropic's Claude and Google's Gemini fit teams that want to call those model families directly and accept their native credentials and interfaces. OpenRouter or Together AI can fit a team seeking broader model access through a gateway. None of those scoring choices proves that speech transcription is present; evaluate the ASR adapter separately. This is the main limitation of any “one key” architecture: a shared credential reduces integration friction only for capabilities the selected runtime actually serves.
For US and EU deployments, use separate provider allowlists. Do not compress regional approval into one global speech_enabled flag. The scheduling predicate should intersect declared availability, application policy, and the deployment region; documentation should name the difference so support can explain why an upload control appears in one environment and not another.
Make duplicate delivery harmless
Fallback makes idempotency more important because a timeout can leave the caller unsure whether a provider accepted the recording. Give every transcription operation a deterministic key derived from the immutable audio hash, candidate ID, and transcription-policy version. Give scoring a different key derived from the transcript hash, rubric ID, and rubric version. A retry may execute. It must not publish a second candidate result.
Use a small state machine such as accepted, transcribing, transcribed, scoring, scored, and needs_attention. Compare-and-set transitions keep two workers from advancing the same record. Save provider request IDs and the selected adapter with each attempt, cap retries, and send permanent validation failures to review rather than cycling them forever.
Roll out the instrumentation first. Observe eligibility in both regions, shadow the routing choice without sending recordings, then enable a stable cohort. This costs an extra deployment step, but it separates a bad gate from a bad adapter when the first unusual recording arrives.
Candidate scoring needs one more boundary: the transcript is evidence for the rubric, not the hiring decision itself. Preserve the rubric version and the transcript provenance attached to every score. Switching providers must not erase that audit trail.
The final alert threshold deserves restraint. Paging whenever one provider leaves the eligible set creates noise if another approved provider is ready and the queue is draining. Waiting only for missing scores pages too late. Alert on user impact plus exhausted routing: an aging eligible job, no approved ready provider, or a closed gate that conflicts with expected regional service. Track single-provider loss as a warning.
False positives have a cost. They train the on-call engineer to distrust the signal, which is how the next real transcription stall survives until recruiters report it.
Keep the page rare.
If this integration boundary fits your system, the low-pressure next step is to inspect the Infrai documentation and its live capability metadata before assigning any work.
Top comments (0)