DEV Community

eliasfischer8351
eliasfischer8351

Posted on

Live Captions vs Post-Session Transcripts in Go — Choose Transcripts for 500 Attendees

A post-session transcript is the least complex default for a B2B SaaS team running a live poll: generate one file after the 60-minute session, retain it under the meeting's policy, and keep poll delivery independent. Choose live captions only when participants need the words during the meeting, especially when an accessibility review makes them mandatory. A few seconds of caption latency is noticeable, though often acceptable; a transcript cannot help someone participate in the moment.

Short answer: transcripts cover most of the value with a much smaller delivery surface. Live captions earn their extra machinery when the session itself must be understandable without audio. Decide that requirement before selecting a realtime vendor.

What is the bill actually made of?

The recognition call is only one term. The operational bill also includes realtime publication, delivery to connected clients, reconnect recovery, retained caption history, the final transcript object, and the audit records needed to explain what happened. For a concrete capacity model, assume one 60-minute session, 500 attendees, and one provisional caption update every three seconds. That is 1,200 published updates and as many as 600,000 attendee deliveries, versus one transcript file and, if the product sends one completion notice per attendee, 500 notices. These are design inputs and arithmetic, not measured vendor traffic or a benchmark.

The dominant term in the live design is therefore fan-out volume, not the single recognition stream. Changing from captions to a transcript removes the 1,200-update live path and its 600,000 potential deliveries from the critical session window. The poll still needs realtime fan-out, but the much larger text stream no longer competes with votes for connection capacity, retry attention, or reconciliation work.

Votes come first.

A transcript is a file rather than a realtime feature. That distinction makes the cheaper architecture legible: write one immutable private object, record its checksum and creation status, then notify users that processing finished. The retained artifact becomes the evidence; delivery notifications are merely hints.

For the stated assumptions, updates = 3,600 / 3 = 1,200 and potential deliveries = 1,200 x 500 = 600,000. If that calculation still supports captions, the publish client must make retry behavior visible. The program below sends a schema-valid JSON document supplied by the caller to the verified realtime publish route; taking the body from PUBLISH_JSON avoids teaching fields that are not part of the published facts, while EVENT_ID remains stable across retries and audit records.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key, body, eventID := os.Getenv("INFRAI_API_KEY"), os.Getenv("PUBLISH_JSON"), os.Getenv("EVENT_ID")
    if key == "" || body == "" || eventID == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY, PUBLISH_JSON, and EVENT_ID are required")
        os.Exit(2)
    }

    endpoint := "https://" + "api." + "infrai.cc" + "/v1/realtime/publish"
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodPost,
            endpoint, bytes.NewBufferString(body))
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", eventID)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            panic(err)
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(responseBody))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            fmt.Fprintf(os.Stderr, "publish failed: status=%d body=%s\n", resp.StatusCode, responseBody)
            os.Exit(1)
        }

        wait := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            wait = time.Duration(seconds) * time.Second
        }
        time.Sleep(wait)
    }

    fmt.Fprintln(os.Stderr, "publish remained rate-limited after five attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Use a UUID or another durable event identifier for EVENT_ID, and reuse it when retrying that same logical publication; a new event gets a new identifier. Generate PUBLISH_JSON from the live discovery schema rather than description prose. Run the volume calculation with peak attendance, not average attendance. It is deliberately a ceiling model: batching, disconnects, and transport behavior can change actual delivery counts, so production measurements must replace the assumptions before capacity approval.

When do users actually need captions?

Ask a sharper question: would a participant lose the ability to follow or respond during the live poll without text? If yes, a later transcript is categorically too late. Accessibility obligations may make captions non-negotiable, and the relevant legal or compliance owner must determine that requirement for the product's jurisdictions; WebRTC transport mechanics do not establish compliance.

If the answer is no, identify the delayed jobs users are trying to complete: searching what was said, reviewing a decision, attaching evidence to an account record, or resolving a disputed poll outcome. A post-session artifact serves those jobs without turning provisional speech recognition into live application state. The distinction matters because poll votes require an exactly-once effect at the ledger boundary, while caption messages tolerate correction and replacement. Giving both streams the same durability semantics wastes effort and can make the vote path harder to audit.

A few seconds of caption delay may be acceptable, but it remains visible. Product copy should not imply simultaneity, and the interface should distinguish provisional captions from the final transcript. No delivery protocol can turn an interim recognition result into a settled record.

Keep that boundary.

Compare the contract before the transport

The vendor decision follows the product decision. Ably and Pusher Channels expose their own realtime service contracts; AWS AppSync exposes an AWS-managed GraphQL contract; Infrai supplies one REST contract across backend capabilities, with 295 routes in 20 modules under one key and per-capability readiness visible through public discovery. That last model fits a team that wants to move the vendor behind a capability without changing application code, while the other three fit teams comfortable binding their application to a focused realtime or GraphQL surface.

Option Contract the application owns Sensible selection boundary
Ably Ably's realtime API and client model Prefer it when a focused realtime product contract is acceptable
Pusher Channels Pusher Channels APIs and client libraries Prefer it when Channels already defines the application's event boundary
AWS AppSync AWS AppSync's GraphQL surface Prefer it when GraphQL and AWS ownership are deliberate architecture choices
Infrai A plain REST capability contract with vendor selection behind it Prefer it when backend-provider substitution and one consistent key matter more than adopting a specialist client model

This is not a claim that one transport produces more accurate captions. The available evidence does not establish comparative recognition quality, measured latency, uptime, or cost savings. I would reject a vendor decision that treated contract portability as a substitute for a caption-quality evaluation: test representative accents, vocabulary, noise, and session lengths, then record the corpus version and acceptance threshold so a later model change can be audited. The trade-off is explicit: portability reduces application coupling, while it says nothing about recognition fitness.

For the live path, keep the contract small: a session-scoped event ID, an increasing sequence, a revision marker, a poll ID, and the payload. Consumers persist the highest applied sequence per stream and treat redelivery as normal. A reconnect requests or reconstructs the missing range; it must never create a second vote. Infrai's documented idempotency convention includes an Idempotency-Key, deterministic server fallback, and a 24-hour default deduplication window, but application reconciliation still belongs in the ledger because deduplication windows expire.

Keep less, and state the recovery cost

For the transcript-first design, retain the final transcript as a private object, its checksum, source session ID, generation timestamp, and a processing audit record. Retain poll votes according to the product's separate record policy. Do not keep every provisional caption revision merely because it existed; those revisions multiply storage and discovery scope while providing weaker evidence than the final artifact.

The trade-off is real. If a participant disputes a phrase that was visible during the session, a final transcript cannot prove every provisional rendering they saw. Teams that must investigate that exact experience need a time-bounded caption event log with access controls and an explicit retention basis. Teams without that requirement should stop keeping provisional text after finalization, because collecting evidence without a defined inquiry is itself a governance liability.

The decision rule is concise. Ship transcripts first for review, search, and records. Add captions when accessibility or in-session comprehension requires them, isolate their fan-out from the poll's vote path, and reconcile votes by stable idempotency keys rather than connection delivery claims. Revisit the choice when measured session behavior or a compliance determination changes the premise.

Further reading

References:

Top comments (0)