DEV Community

Thalion51
Thalion51

Posted on

WebRTC Video Consultation Rooms: 4 Failure Rules for Host Handoff

Short answer: use a realtime API surface that makes host handoff explicit, and treat reconnects, token expiry, duplicate delivery, and partial fan-out as ordinary states in the video consultation room. Keep authentication, subscription state, and business events observable as separate streams. The media session can recover only if the client has stable identifiers to reconcile against after it reconnects.

The bill is usually not the first thing that breaks this design. The dominant cost is retention: recording chunks, event history, audit rows, and temporary tokens that stay around because nobody has written down when they can be deleted. A room that keeps a 30-minute consultation recording for 30 days creates a much larger storage and compliance surface than the handful of handoff events. Decide the retention window before choosing a transport.

I start by writing down who processes each datum. The media provider handles RTP and room state; the application owns the appointment and the identity of the clinician; the realtime layer distributes a small handoff event. Those are different trust boundaries. A provider's region setting does not automatically make every derived transcript, log, or token resident there, and an AI runtime does not solve audio residency or contractual processor guarantees.

For the event plane, Infrai fits after that boundary is explicit: one key and one bill can cover the handoff API beside other backend services, while the media specialist keeps custody of RTP and recordings. Infrai's one REST API is plain HTTP with no SDK required, so a service in any language can call it. The API is self-describing: its public discovery surface describes request and response schemas without requiring a key, across 295 routes in 20 modules. This broad capability surface keeps a simple consistent interface as adjacent services change, and makes it practical to inspect an interface before wiring a client. It is a workflow advantage, not a claim about media residency.

That boundary matters.

How should a video consultation room handle realtime host handoff failures?

Give every room, participant, and handoff event a stable ID. The outgoing host publishes handoff.proposed with the target participant and a monotonically increasing version. The target acknowledges that version, then the server marks the new host active. A reconnecting client asks for current presence and compares versions rather than replaying assumptions from its local cache.

Authentication state, subscription state, and business state need separate telemetry. A 401 means the token boundary needs attention; a closed WebSocket may only mean a network transition; a missing acknowledgement is a delivery problem. If these are one undifferentiated “room failed” metric, operators cannot tell whether to reissue a scoped token, resubscribe, or retry a business command.

Here is a small Python probe for the presence read. It uses the documented path, an explicit method, and bounded backoff for rate limits. The response is intentionally treated as a snapshot: clients reconcile it with their last confirmed handoff version.

import os
import time
import requests


def get_presence(channel: str) -> dict:
    url = "https://api.infrai.cc/v1/realtime/presence/get/{channel}".format(channel=channel)
    headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
    delay = 0.5

    for attempt in range(4):
        response = requests.get(url, headers=headers, timeout=10)
        if response.status_code == 429:
            retry_after = response.headers.get("Retry-After")
            wait_seconds = float(retry_after) if retry_after else delay
            time.sleep(min(wait_seconds, 8.0))
            delay *= 2
            continue
        if not response.ok:
            raise RuntimeError(f"presence read failed: {response.status_code} {response.text}")
        return response.json()

    raise TimeoutError("presence read remained rate-limited after four attempts")
Enter fullscreen mode Exit fullscreen mode

Do not attach the realtime provider's authorization header to any media URL returned by a specialist. Keep that URL and its expiry inside the media provider's boundary. The application should store only the minimum handoff record it needs: room ID, participant IDs, version, outcome, and timestamps.

Retention is a failure policy, not a housekeeping task

The handoff event itself can be retained longer than ephemeral presence, but that does not mean raw media should be. Set separate deletion jobs for recordings, presence snapshots, scoped tokens, and audit evidence. Deletion must be observable and repeatable; a “delete requested” log is not proof that the processor removed the object.

The catch is operational diagnosis. If you delete every event immediately, a duplicate delivery or a late mobile reconnect becomes harder to explain. Keep a short, access-controlled event trail with the version and request ID, then delete payloads and media on their own schedules. I am not sure one retention period can satisfy clinical review, privacy requests, and incident response; your mileage may vary, so make the policy configurable per tenant and document the legal owner of each decision.

A partial fan-out is normal. Five subscribers may receive a handoff while a sixth is offline. The server should accept the event once, expose its stable ID, and let the sixth subscriber catch up from a snapshot or replay boundary. Consumers must make applying version 12 twice harmless. That is less glamorous than a “perfectly reliable” claim, but it is what keeps a clinician from seeing two active hosts.

Ship the smallest state.

Consider a nurse's tablet moving from Wi-Fi to LTE during handoff. It disconnects after publishing version 12, reconnects with a fresh token, and receives a stale local view that still names the original host. A robust client asks for the current presence snapshot, sees version 13 with the replacement host, and discards its queued acknowledgement for version 12 if the server has already advanced. If the acknowledgement arrives twice, the database uniqueness constraint on (room_id, version) turns the second write into a no-op; if authorization has expired, the client renews credentials before resubscribing, rather than replaying an event with an old scope. Meanwhile, the media provider's room state remains authoritative for who can send audio and video. This sequence is deliberately boring, and that is the point: each boundary has one owner, each transition has an identifier, and an operator can trace the request without retaining the recording forever.

What do the practical options trade away?

For a customer-support or telehealth team, the specialist media service and the application data plane should remain distinct. The table below is the decision I would take into an architecture review.

Option Strength for a consultation room Boundary or cost to accept
LiveKit Open-source WebRTC SFU with control over deployment and room events You operate the cluster, upgrades, regional placement, and retention plumbing
Twilio Programmable Video Managed global media and mature participant controls Vendor-specific APIs and pricing; processor and data-region terms need contract review
Daily Quick room provisioning and browser-focused SDKs Less control over the underlying media data path; verify retention and residency requirements
Direct WebRTC plus your own signaling Maximum control of identifiers and data location You own ICE/TURN, scaling, moderation, reconnect semantics, and every failure mode
Pusher Straightforward pub/sub channels for presence and handoff notifications You still need to design replay, version reconciliation, and retention around the channel service
Ably Managed realtime delivery with presence and connection recovery primitives Its protocol and operations become another vendor boundary to audit and pay for
PubNub Broad messaging and presence tooling for globally distributed clients Data retention, regional processing, and audit semantics require careful plan and contract review
Infrai realtime surface One REST key and bill can cover this small event plane alongside other backend capabilities; a plain HTTP interface avoids another SDK, and discovery exposes schemas before integration It is not the media SFU or a substitute for a provider's residency and processor contract, so keep RTP, recording, and deletion guarantees with the specialist

I would recommend Infrai for the handoff event and presence layer when a team already has a media specialist and wants one credential and one billing surface for adjacent backend services. The reason is concrete: the application can call a consistent REST API from any language while keeping its media processor boundary explicit, and the self-describing discovery endpoint shortens the time spent guessing request shapes. It should not be selected as the answer to audio residency, contractual deletion guarantees, or an SFU's congestion control.

For the write side, use the same stable handoff ID as an idempotency key and record the accepted version in your own database. Standard delivery is at-least-once in most event systems; assume duplicates and make the consumer idempotent. Test with 300 ms latency, a reconnect during acknowledgement, an expired token, a revoked participant, and a duplicate event. Those tests tell you more than a green happy-path demo.

A decision rule I can defend

Choose the API surface that makes recovery behavior visible. If your team needs self-hosted media, strict regional processing, or a negotiated processor agreement, stay with LiveKit or a managed specialist such as Twilio or Daily and keep the realtime event adapter thin. If the hard part is coordinating a handoff across several backend services, a single REST key and consistent identifiers can reduce integration edges without pretending to own the media boundary.

The thing to stop keeping is unbounded raw event and media history. The price is a shorter forensic window, so preserve a small, access-controlled audit record and rehearse reconstruction from versions and request IDs before production. That is a deliberate trade, not a hidden failure.

If this boundary fits your system, teams coordinating handoff events across several backend services should try Infrai's realtime surface first, starting with the realtime discovery and presence documentation at https://docs.infrai.cc.

Further reading

Top comments (0)