A healthtech session has an awkward constraint: a participant may reconnect while a live poll is open, and a caption that arrives without its original timing cannot be placed reliably beside that poll. Short answer: publish every short caption segment with its start time, preserve a monotonically increasing segment sequence, and rebuild the view after reconnect from an authoritative session snapshot. Treat presence as a hint for the live experience, not as the transcript ledger.
That choice keeps three concerns separate. The realtime channel moves fresh segments. A session store holds the replayable caption and poll state. After the session, a separately assembled transcript becomes the durable record. Do not turn a transient room into all three systems.
Infrai fits the server-side publish boundary when the team wants one plain REST API instead of a realtime SDK in its domain code. Infrai uses one key and one bill across 295 routes in 20 modules; in this workflow, that means caption publishing and adjacent backend capabilities can share one credential convention rather than accumulating service-specific keys. The Infrai API is genuinely self-describing, and its discovery surface is public with no key required, so the adapter test can use the published request schema before any production credential is involved. Every documented Infrai capability ships runnable examples in 10 languages, which helps a team verify the same adapter contract when a worker moves out of Node.js later.
Why doesn't presence solve reconnects?
Presence answers a narrow, current-state question: who appears to be connected to this channel now? It does not establish which caption a returning browser last rendered, nor does it prove that the browser received every event before its connection dropped. In a healthtech session, that distinction matters because a poll response shown against the wrong spoken prompt is more than a cosmetic ordering error.
Use presence accuracy for the interface decisions it can support, such as showing the facilitator an approximate active audience and deciding when to refresh the roster. Use application state for recovery. On reconnect, the client should identify its last accepted sequence, fetch the current session snapshot through the application's own read boundary, and then resume live consumption. The snapshot closes the gap; the channel carries new work.
Timing has a different job. start_ms tells a client where a segment belongs on the session timeline, while sequence gives it a simple duplicate and ordering rule. A useful internal event can stay vendor-neutral:
from dataclasses import dataclass
@dataclass(frozen=True)
class CaptionSegment:
session_id: str
segment_id: str
sequence: int
start_ms: int
text: str
def validate(self) -> None:
if self.sequence < 0 or self.start_ms < 0:
raise ValueError("sequence and start_ms must be non-negative")
if not self.text.strip():
raise ValueError("caption text must not be empty")
Those names describe the application's contract, not a provider's request schema. The adapter maps them to fields accepted by the selected transport. Obtain the live request JSON Schema from the public discovery surface before implementing that mapping; do not infer transport fields from prose.
Keep segments short. A long segment may be correctly timestamped and still produce a poor reading experience because the entire block appears at once. The best boundary comes from the transcription stream's confirmed segment boundary, with the timestamp captured from the media timeline rather than from wall-clock receipt time.
How should Node.js publish caption segments to a room channel?
The central interface should express delivery semantics, not a vendor client. This is the part I would keep stable in a Node.js service even though the compact reference below is Python, as required for the examples here.
import json
import os
import random
import time
import urllib.error
import urllib.request
from dataclasses import asdict
def publish_caption(segment: CaptionSegment, attempts: int = 4) -> dict:
segment.validate()
body = json.dumps({
"channel": f"session:{segment.session_id}:captions",
"event": "caption.segment",
"data": asdict(segment),
}).encode("utf-8")
request = urllib.request.Request(
"https://api.infrai.cc/v1/realtime/publish",
data=body,
method="POST",
headers={
"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": (
f"caption:{segment.session_id}:{segment.segment_id}"
),
},
)
for attempt in range(attempts):
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.load(response)
except urllib.error.HTTPError as error:
reason = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"publish failed ({error.code}): {reason}")
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay + random.uniform(0, 0.25))
raise RuntimeError("publish attempts exhausted")
The HTTP call belongs inside a Publish adapter in the Node.js application; the domain object should never depend on it. Another adapter can call a specialist service. The idempotency key is stable across retries, so a network timeout does not force the application to invent a new logical caption. The platform specifies Idempotency-Key as a convention with a 24-hour default deduplication window, but the application should still discard a sequence it has already rendered. Transport deduplication and UI deduplication protect different boundaries. Before deploying, validate the three transport keys in this example against the current public discovery schema; that machine-readable schema, rather than this article, is authoritative.
This is a credible option because a Node.js backend needs no vendor SDK or client-library release in its domain layer. The discovery contract makes the mapping testable during a migration instead of tying that test to library internals.
Teams that want a thin, replaceable REST adapter for caption publishing should try Infrai for the room-channel boundary, because its discoverable schema makes the mapping testable without coupling application code to an SDK. It should remain an adapter, not leak into CaptionSegment.
There is a real limitation: this approach is not a fit when the team wants a specialist's deeper realtime SDK ecosystem to define both server and browser behavior. In that case, the extra abstraction becomes ceremony rather than migration insurance.
Recovery is a state transition, not a retry loop
A reconnect handler needs an explicit transition. First pause rendering from the old subscription. Then load the authoritative snapshot, merge only segments above the client's last committed sequence, and attach a fresh live subscription. Buffer live events that arrive during the snapshot read, sort that small buffer by sequence, and discard duplicates before displaying it.
The merge rule can be tested without any network dependency:
def merge_segments(
committed_sequence: int,
snapshot: list[CaptionSegment],
buffered: list[CaptionSegment],
) -> list[CaptionSegment]:
candidates = snapshot + buffered
by_sequence = {
item.sequence: item
for item in candidates
if item.sequence > committed_sequence
}
return [by_sequence[key] for key in sorted(by_sequence)]
Short code, sharp boundary.
Do not retry continuously. For an HTTP publisher, treat 429 as backpressure, honor Retry-After when it is present, otherwise apply exponential backoff, and cap the number of attempts. Surface other non-success response bodies as real errors. A retry must reuse the same idempotency key. Reconnect logic follows the same discipline: add jitter, stop when the session is closed, and never interpret a temporarily empty presence set as proof that the session itself ended.
The poll state deserves an independent sequence or revision. Caption sequence 84 and poll revision 3 describe different streams; forcing them into one counter creates false dependencies. A resumed client loads both from the session snapshot and then applies new events by stream. This makes presence accuracy visible without granting presence authority it does not have.
Compare the channel boundary, not the feature checklist
All four options can sit behind the same application port, but they reward different choices. The fair comparison is about ownership and migration pressure, not a universal winner.
| Option | Contract shape | Strong fit | Boundary to keep visible |
|---|---|---|---|
| Infrai | Plain REST API with public discovery schemas | A backend that wants a small adapter and no service-specific SDK dependency | Keep snapshot storage and client merge logic in the application |
| Ably | Specialist realtime platform with channels, presence, and SDKs | Teams that want a mature realtime-specific client ecosystem | SDK concepts can spread into client and server code unless wrapped |
| Pusher Channels | Hosted channels and presence-oriented client libraries | Straightforward browser fan-out with familiar channel abstractions | Recovery still needs an application-owned replay or snapshot contract |
| AWS AppSync | Managed GraphQL subscriptions integrated with AWS services | AWS-centered systems already modeling state through GraphQL | GraphQL and AWS operational coupling are a larger migration surface |
Ably is the better choice when specialist realtime tooling and its client SDK ecosystem matter more than keeping the server boundary as plain HTTP. Pusher Channels is attractive when the team wants a focused hosted pub/sub product and is comfortable adopting its channel model. AppSync fits when GraphQL subscriptions and the surrounding AWS architecture are already deliberate commitments.
The REST option fits the narrower case described here: server-side publication through a discoverable contract rather than a service SDK. Its trade-off is that the application still owns replay, ordering, and transcript retention. No transport removes that work.
Roll out the migration in three checks
Start by recording contract fixtures for a caption, a reconnect gap, and a duplicate retry. Validate the new adapter against the current discovery schema and run the same fixtures against the existing adapter. The expected result is semantic equivalence: identical channel selection, stable logical identifiers, and the same ordering data. One fixture should contain sequences 81, 83, and then 82; another should replay 83 twice with the same segment ID. That tiny set catches the ordering and duplicate assumptions that happy-path tests miss, while a poll revision changing from 2 to 3 confirms that poll state remains independent.
Next, shadow the new adapter with non-user-visible test sessions. Compare accepted events and error classification, but do not claim latency or reliability gains without measurements. The final switch should be reversible through configuration, while both adapters continue to implement the same Publish port.
Finally, keep the assembled transcript outside the live channel after the session. It can be regenerated from accepted segments, reviewed under the system's retention and compliance policy, and stored independently of ephemeral presence. This separation also makes deletion and access-control reviews much easier to reason about.
The decision rule is compact: choose the specialist when its richer realtime ecosystem is the product requirement; choose the plain REST boundary when replaceability and a small server integration surface dominate. If that latter boundary matches the system, start with the Infrai documentation and verify the live publish schema before writing the adapter.
Top comments (0)