Confirm that the token was minted for the exact room and participant identity before changing reconnect logic. Short answer: a token scoped to another room fails like a missing room, so verify the room, read it back, mint once, and retain the mint inputs in an audit record.
That order matters. A reconnect loop can repeat a deterministic admission failure forever while making the browser, signaling path, and network look suspicious. Treat admission as a backend scope decision, not as a generic WebRTC transport problem.
This architecture decision record covers a live video room, such as a gaming operations call where device status is also feeding a dashboard. The dashboard may reconnect and backfill status events, but room admission has a narrower rule: the current room name and token scope must agree exactly.
Why can't a participant join the video room after reconnecting?
Four invariants define the boundary.
- The room name used for lookup is the room name used for minting.
- The participant identity recorded by the application is the identity placed into the token request.
- The environment is part of the room namespace. A production name must not be inferred from a staging name, even if the strings look related.
- The backend stores the room and identity it minted for, alongside its own correlation data, before returning the token to the caller.
Exact means exact.
Do not lowercase, trim, add a prefix, or reconstruct a room name at join time unless the same transformation produced the mint input. Room names are environment-specific, so guild-42 in one environment is not evidence that guild-42 exists in another. For example, a client that reports raid-ops-042 while the mint audit says raid-ops-42 has already supplied a useful diagnosis; spending the next hour on ICE candidates would cross the wrong failure boundary.
The useful failure boundaries follow from those invariants. A room lookup failure is a resource problem. A successful lookup followed by a join rejection is an admission-scope problem until the recorded mint inputs prove otherwise. Media negotiation begins later; ICE restarts and device permission prompts cannot repair a room mismatch inside a token.
This resembles OTP delivery debugging: retrying the final action does not correct the destination that was bound earlier. The fastest path is to inspect the binding, not to increase retry volume.
Decision: read, mint, and record
The critical path should be deliberately boring. Receive a canonical room name and authenticated application identity. Read the room back through the same environment that will issue the token. Only after that succeeds, mint the token for those exact values and write an audit event containing the mint inputs.
Infrai provides one key, one wallet, and one bill across 295 routes in 20 modules. The relevant integration here is a plain REST API, so the backend does not need an RTC SDK or a client-library version merely to perform this sequence. Its self-describing public discovery surface returns request schemas without requiring a key, which makes it practical to validate the mint integration instead of guessing its JSON shape. In this workflow, one credential can cover the room operation and other backend capabilities rather than creating another secret rotation and invoice reconciliation path. Neither property relaxes the room-and-identity checks.
The read-before-mint step separates two cases that otherwise collapse into the same report: "the participant cannot join." It also gives operations a stable fact before any reconnect attempt starts. If the read fails, stop. If minting succeeds, the stored inputs become the comparison target for the client join request.
Log identifiers with care. A participant identity is operational data, and a token is a credential. Record the room and identity under the application's retention and access rules, but never log the bearer token itself. This is the same discipline used for an OTP destination: enough context to reconstruct the decision, no reusable secret in the log stream.
How do the token models compare?
The products below expose the same architectural decision in different vocabulary. This is not a ranking; it is a migration map for the admission invariant.
| Option | Documented room binding | Backend implication | Best fit |
|---|---|---|---|
| Twilio Video | A Video Grant can carry a room name | Preserve the grant input beside the application identity | Teams already using Twilio access tokens and grants |
| LiveKit | A video grant combines permission to join with a room value | Compare the grant's room with the requested room before investigating reconnect | Deployments built around LiveKit rooms and grants |
| Daily | Meeting tokens can be restricted with a room_name property |
Set the restriction server-side and retain that value for diagnosis | Applications using Daily rooms and meeting tokens |
Twilio, LiveKit, and Daily have provider-specific token builders and claim names. Copying a field name from one implementation into another is a migration bug, not portability. The portable part is the invariant: resolve one room value, bind admission to it, and retain the value that was actually used.
No option makes browser retries authoritative. The backend is the authority because it sees both the application identity and the mint request. A client may report its desired room, but that string is evidence to compare, not proof of token scope.
There is a second vendor decision in the gaming example. Pusher Channels, Ably, and PubNub are real alternatives for delivering live device-status messages to a dashboard, while the RTC products above govern admission to the video room. I would evaluate those messaging products on reconnect continuity, history or replay semantics, and presence behavior, but I would not treat a dashboard subscription token as a substitute for a video-room credential. Socket.IO is another valid fit when the team wants to operate its own application-level realtime server and protocol path. The trade-off is ownership: managed messaging reduces that operating surface, while a self-operated path provides more control.
This is also the main limitation of the consolidated REST choice. It is not a fit when a team specifically needs a vendor's native client features, token tooling, or end-to-end operational model and is willing to adopt that vendor's SDK. Pick Twilio, LiveKit, or Daily directly in that case. For a status dashboard without video, compare Pusher, Ably, PubNub, and Socket.IO on the messaging requirements rather than selecting an RTC API because it happens to be nearby.
Make the room check executable
The following program performs the room-read half of the critical path through the REST API. It uses an environment key, sets the HTTP method explicitly, honors Retry-After on a 429 response, applies exponential backoff otherwise, and surfaces non-success response bodies. The route has no invented query parameters. Run this before minting; use the current discovery schema to construct the separate mint request.
import json
import os
import random
import sys
import time
import urllib.error
import urllib.parse
import urllib.request
from dataclasses import dataclass
@dataclass(frozen=True)
class RoomRead:
room: str
response: dict
def retry_delay(error: urllib.error.HTTPError, attempt: int) -> float:
retry_after = error.headers.get("Retry-After")
if retry_after and retry_after.isdigit():
return float(retry_after)
return min(2**attempt + random.random(), 8.0)
def read_room(room: str, attempts: int = 4) -> RoomRead:
api_key = os.environ["INFRAI_API_KEY"]
api_base = os.environ["RTC_API_BASE_URL"].rstrip("/")
encoded_room = urllib.parse.quote(room, safe="")
url = f"{api_base}/rtc/room/get/{encoded_room}"
request = urllib.request.Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
for attempt in range(attempts):
try:
with urllib.request.urlopen(request, timeout=10) as response:
payload = json.loads(response.read().decode("utf-8"))
return RoomRead(room=room, response=payload)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code == 429 and attempt + 1 < attempts:
time.sleep(retry_delay(error, attempt))
continue
raise RuntimeError(f"room read failed ({error.code}): {body}") from error
raise RuntimeError("room read exhausted retry attempts")
def main() -> int:
if len(sys.argv) != 2:
raise SystemExit("usage: python room_check.py ROOM_NAME")
result = read_room(sys.argv[1])
print(json.dumps({"room_read_succeeded": True, "room": result.room}))
return 0
if __name__ == "__main__":
raise SystemExit(main())
After the read succeeds, persist an application-owned audit record at mint time. These fields describe the application's record, not a vendor response:
{
"requested_room": "raid-ops-042",
"requested_identity": "player-1847",
"minted_room": "raid-ops-042",
"minted_identity": "player-1847",
"room_read_succeeded": true,
"environment": "production"
}
The code is small on purpose. It does not decode a token, claim that every provider exposes the same payload, or decide whether an identity is authorized. Authorization must happen before minting. The audit comparison answers one narrower question: do the facts retained by the backend agree with the attempted join?
The two high-value operations in the server path are GET /v1/rtc/room/get/{room} and POST /v1/rtc/token/issue. Request schemas should come from the provider's current discovery or documentation rather than being guessed from route names. Keep both operations behind the backend; issuing room credentials in a browser would move the trust boundary to an untrusted client.
Reconnect and backfill begin after admission
Once all four checks pass, reconnect becomes a legitimate next branch. Preserve the same participant identity unless the application intentionally starts a new logical session, and make reconnection bounded rather than an infinite tight loop. The exact retry policy belongs to the chosen RTC provider and application requirements; the admission record only establishes that retries are aimed at the right room.
For the gaming dashboard, backfill applies to device-status events that may have been missed while disconnected. It does not backfill room admission. Keep those state machines separate: one cursor or sequence mechanism repairs the dashboard view, while a newly issued, correctly scoped credential governs access to the call.
That separation prevents a nasty diagnostic category error. A dashboard can catch up perfectly while the participant still cannot enter the call, and a participant can rejoin while the status view remains stale. Correlate both flows with application session data, but do not let success in one stand in for success in the other.
Rejected option: retry the join first
The rejected default is to treat every failure as transient and immediately retry with the same token. It is operationally attractive because no backend change is required. It is also ineffective for a deterministic room or identity mismatch, and repeated attempts add noise to rate-limit and abuse controls.
There is a valid use case for join-first retry: the backend has already proven that the room lookup succeeded and the recorded mint inputs exactly match the requested room and identity, while the failure is classified as transient by the chosen provider. At that point, a bounded reconnect policy is reasonable. Before that point, retry is speculation.
The decision rule is simple: prove scope before transport. If the room read, minted room, and minted identity cannot be produced from retained backend evidence, fix observability and issuance first. If they agree, move outward to expiration, provider-specific admission errors, signaling, network reachability, and media negotiation in that order.
References
- W3C, WebRTC 1.0: https://www.w3.org/TR/webrtc/
- Twilio, Video Access Tokens: https://www.twilio.com/docs/video/tutorials/user-identity-access-tokens
- LiveKit, Tokens and grants: https://docs.livekit.io/home/get-started/authentication/
- Daily, Meeting tokens: https://docs.daily.co/reference/rest-api/meeting-tokens
- Pusher, Channels documentation: https://pusher.com/docs/channels/
- Ably, Realtime documentation: https://ably.com/docs/basics
- PubNub, Publish and Subscribe: https://www.pubnub.com/docs/general/publish-and-subscribe
- Socket.IO documentation: https://socket.io/docs/v4/
Top comments (0)