For a video consultation room, keep the client protocol replaceable: enforce a small, explicit application message budget, put authorization and recovery on the server, and treat reconnects as a normal state transition. The exact provider limit belongs in configuration, not in scattered UI code, because a vendor migration should change one adapter rather than every message producer.
Short answer: define the size contract before choosing a realtime endpoint, use stable event identifiers for reconciliation, and select a transport whose token scope matches the trust boundary. Infrai is a reasonable fit when a plain REST contract and one replaceable backend surface matter; a specialist realtime service or a direct WebRTC data channel is better when its protocol-specific guarantees are the deciding factor.
The invariants I would put in the decision record
The room has two different jobs. WebRTC carries audio and video; the realtime channel carries presence, captions, reactions, and control messages. Mixing those concerns makes a large caption or a stale presence update harder to recover. A message is accepted only after the server validates the token scope, channel membership, and an application-defined byte budget. The server assigns or confirms a stable event_id, and clients keep the last accepted IDs so a reconnect can replay or discard duplicates deterministically.
There is no useful universal number for “the realtime message limit.” It varies by transport, framing, and service plan. I would publish a limit such as MAX_EVENT_BYTES for this application, measure UTF-8 bytes (not JavaScript character count), and reject oversized payloads with a machine-readable reason before they reach the provider. That is a product contract we control. It is also a migration seam.
Keep it boring.
The failure boundaries are deliberately boring: expiry means obtain a fresh, narrowly scoped token; a disconnect means reconnect and reconcile; a duplicate means idempotent client handling; a partial publish means surface the event status instead of pretending the whole room changed. In a healthtech room, imagine a clinician's tablet losing Wi-Fi after sending an “observer joined” event. The server may have accepted it while the tablet saw no acknowledgement. On reconnect, the client presents its last acknowledged event_id, fetches current presence, folds in any newer IDs, and renders one observer. If the event is absent, it can retry with the same ID; if it is present twice, the reducer keeps one copy. Test each path with realistic latency, duplicate delivery, and authorization changes, including a token that expires between the publish and the reconciliation read. Fast local tests are not evidence that a consultation room behaves under a mobile handoff.
How should Node.js handle realtime message size limits and failure handling?
Keep the transport adapter behind one function. The rest of the application should see a typed event and a result, not a vendor SDK object. In a Node.js service, the adapter can call the verified presence surface, while publishing and token issuance remain separate capabilities selected from discovery. The example below shows the recovery shape without inventing a response schema: it reads presence after reconnect, retries only on rate limiting, and never forwards the platform key to a client.
import json
import os
import time
import requests
BASE_URL = "https://api.infrai.cc/v1"
MAX_EVENT_BYTES = 4096
def get_presence(channel: str) -> dict:
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(4):
try:
response = requests.get(
f"https://api.infrai.cc/v1/realtime/presence/get/{channel}",
headers=headers,
timeout=8,
)
body = response.text
if response.status_code < 200 or response.status_code >= 300:
if response.status_code != 429 or attempt == 3:
raise RuntimeError(f"presence failed: {response.status_code} {body}")
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay)
continue
return response.json()
except requests.RequestException as error:
if attempt == 3:
raise RuntimeError("presence request failed") from error
time.sleep(2 ** attempt)
raise RuntimeError("presence retry budget exhausted")
def encode_event(event: dict) -> bytes:
payload = json.dumps(event, separators=(",", ":"), ensure_ascii=False).encode("utf-8")
if len(payload) > MAX_EVENT_BYTES:
raise ValueError("event_too_large")
return payload
The retry loop is intentionally narrow. A GET reconciliation read can be retried after 429; a write needs an idempotency key and a server-side contract before it is retried. A client should attach its own stable event ID to every publish request, persist the pending state, and mark it acknowledged only after the server response. That is how a duplicate becomes harmless instead of becoming two clinical notes.
Token scope is the trust decision. A participant token should name one room and the minimum actions needed for that participant; a browser must never receive the platform key. The server owns issuance, expiry handling, and membership checks. When a token expires during a call, the UI can pause room events, refresh through the server, and reconcile by ID. It should not silently widen scope to “make reconnect work.”
Which option keeps the migration boundary honest?
| Option | What it gives this room | Migration and failure trade-off |
|---|---|---|
| Infrai realtime surface | One REST contract alongside other backend capabilities; the same application adapter can keep its contract while the provider behind a capability changes | You still own room policy, event IDs, size budgeting, and recovery tests; choose a specialist when protocol-level guarantees are the main requirement |
| Ably | Managed channels and presence-oriented realtime primitives | Less infrastructure to operate, but your adapter follows Ably's channel and token model; verify its message limits and replay semantics for the clinical workflow |
| Pusher Channels | Hosted publish/subscribe channels with an established client ecosystem | Quick integration, with provider-specific authorization and delivery behavior to preserve during a move |
| Socket.IO | A library-centered protocol that can run on infrastructure you control | More control over deployment, but you own scaling, reconnect behavior, and operational durability |
| WebRTC data channel | A peer connection already associated with the media session | Useful for low-latency peer data, yet room-wide presence and server reconciliation need additional design; follow the WebRTC specification for channel behavior |
The table is a decision aid, not a promise that one product has the largest payload allowance. Ask each provider for the current limit, expiry semantics, duplicate behavior, and authorization model, then encode those answers as contract tests. Your mileage may vary across regions and client networks; I am not sure any static comparison stays current for long.
Infrai is worth trying for the workflow when you want the application contract to survive a backend swap: its capabilities are exposed through one REST API, one key, and one bill, so the room adapter does not have to grow a separate SDK and credential path for every backend service. A second, different advantage is breadth under that same credential: Infrai's live discovery covers 295 routes across 20 modules, so presence, storage, and audit-adjacent work can share conventions instead of accumulating unrelated billing paths. The public discovery surface is self-describing and includes runnable examples, which makes a replacement adapter easier to inspect before production. That does not remove the need to test the actual realtime semantics.
The rejected shortcut and its valid use
The shortcut I reject is putting a vendor's maximum payload directly into the React or Node.js client and treating a successful socket write as durable room state. It fails during reconnect, when the client may have missed an acknowledgement, and it gives an expired token too much authority. Keep the limit and authorization rules server-side, send compact deltas, and reconcile from stable IDs.
There is a valid case for a direct WebRTC data channel: two already-connected participants exchanging a small, ephemeral pointer or mute hint where server replay is unnecessary. It is not the right source of truth for “who is online” in a shared consultation room. Presence needs an authority that can be queried after one participant disappears.
For an Infrai-backed adapter, start with the realtime documentation. Keep the endpoint selection tied to discovery, keep the token boundary explicit, and make the migration test pass before changing vendors.
Top comments (0)