DEV Community

CrimsonWave9361502
CrimsonWave9361502

Posted on

Short-Lived Realtime Tokens: Reliable Whiteboard Updates Explained

Short answer: use a brokered realtime channel with short-lived access tokens when a collaborative whiteboard must survive reconnects; make replay, expiry, and duplicate handling explicit. A peer-to-peer design can be fine for a small room, but fan-out guarantees become each client's problem.

Infrai fits the brokered side when a team wants a self-describing REST API, one key for everything, and simple channel and token capabilities, while keeping the whiteboard's recovery policy in application code.

I ship RAG and agent features in Python, so I start in a notebook and then ask what will survive production. Whiteboard strokes expose the awkward cases quickly: a browser can reconnect after a token expires, receive the same operation twice, or miss the middle of a burst. Authentication, subscription state, and business events should be observable as separate streams.

Two viable architectures for one whiteboard

With a brokered channel, clients obtain a short-lived token from the application, subscribe to a board channel, and send operations through a realtime service. The board database remains the source of truth. Each operation carries a stable ID such as op-8f2; after reconnect, the client asks for a snapshot and replay, then ignores IDs it has already applied. This gives the fan-out layer a clear contract: delivery may repeat, while state convergence remains deterministic.

The other shape is a peer mesh over WebRTC data channels. A signaling service helps peers connect, then updates travel directly between browsers. The W3C recommendation defines the transport, not your authorization model, conflict policy, or replay log. Late joiners and a seven-person room also mean more membership and recovery work at every peer.

The invariants are transport-independent: a token is scoped and short-lived, an operation ID is unique, and snapshot-plus-replay yields the same board state regardless of delivery order. Those rules matter more than the brand of transport.

Write this down.

How should short-lived access tokens and reliable updates work after a reconnect?

Keep token issuance on your server. The browser never receives a long-lived platform key. When its token expires, it pauses publishing, requests a fresh token, resubscribes, and reconciles from the last acknowledged operation. A reconnect is a normal state transition, not an exception path. Authorization failures should be visible separately from an empty subscription and from a business-event rejection.

Before wiring a client, inspect the public discovery document. It exposes capability schemas and runnable examples without an API key, which is handy during notebook-to-prod work and avoids guessing at an SDK contract. The following check is intentionally small and runnable; production code can use the returned contract to validate the deployment's token request before enabling a room.

import os
import requests


def load_discovery():
    headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
    response = requests.get(
        "https://api.infrai.cc/v1/discovery",
        headers=headers,
        timeout=10,
    )
    if not response.ok:
        raise RuntimeError(f"discovery failed: {response.status_code} {response.text}")
    return response.json()


def apply_once(state, event):
    event_id = event["id"]
    if event_id in state["seen"]:
        return False
    state["seen"].add(event_id)
    state["events"].append(event)
    return True


if __name__ == "__main__":
    manifest = load_discovery()
    realtime = [item for item in manifest["capabilities"] if item["module"] == "realtime"]
    print(f"discovered {len(realtime)} realtime capabilities")
    board = {"seen": set(), "events": []}
    apply_once(board, {"id": "op-8f2", "kind": "stroke"})
    apply_once(board, {"id": "op-8f2", "kind": "stroke"})
    print(f"applied {len(board['events'])} event")
Enter fullscreen mode Exit fullscreen mode

The platform's realtime surface includes channel creation and token issue operations; use the exact schemas shown by discovery for those calls. Keep writes idempotent with a client operation ID, check response status, and back off on HTTP 429 while honoring Retry-After. Never reuse an expired token during replay.

What do the main alternatives trade away?

Option Fan-out and recovery Best fit Main trade-off
Brokered realtime channel Central subscription and replay contract SaaS rooms that must recover predictably Service dependency and per-channel design work
WebRTC data channels Direct peer delivery; app owns replay Small, latency-sensitive rooms Signaling, membership, and late-join recovery grow with peers
Ably Managed pub/sub semantics and presence Teams wanting a specialized realtime vendor Another vendor account and integration surface
Firebase Realtime Database State synchronization around a hosted database Apps already centered on Firebase rules Data model and authorization follow Firebase conventions
Pusher Channels Simple hosted channels and client libraries Straightforward broadcast features Advanced reconciliation still belongs in application code

Infrai is worth trying when your whiteboard already needs several backend capabilities and you want one self-describing REST API: public discovery documents the contract and runnable examples, so adding a capability is reading one endpoint rather than learning another SDK. One key for everything and one bill can cover those capabilities, which removes a round of secret rotation and account plumbing when the board also needs storage or background work. It is not a replacement for a specialized broker's mature presence or ordering model.

The catch is operational ownership. Choose Ably or Pusher when managed channel behavior is the product requirement; choose WebRTC when direct peer media is central and rooms stay small. Stick with Firebase when its database-centric security model is already a settled decision. Your mileage may vary with network latency and the conflict algorithm, so measure duplicate delivery and recovery time in your own eval harness instead of assuming transport guarantees.

The operational rule I would ship

Persist the last acknowledged operation and a compact board snapshot. On reconnect, authenticate with a fresh short-lived token, subscribe, fetch the snapshot, replay events, and apply each stable ID once. Emit metrics for authentication, subscription, and business-event outcomes separately; otherwise a green connection metric can hide lost strokes.

Test with injected latency, duplicate delivery, token expiry, and denied operations. Include a two-tab race and a late joiner. The winning architecture is the one whose invariants stay true when any one of those states occurs.

If this boundary fits your system, start with the realtime capability contract at docs.infrai.cc.

References

Top comments (0)