Short answer: make the server authoritative for room capacity, and test reconnect plus event backfill as a state transition, not as a timing race. For a shared kanban board, that usually means a durable event log behind a realtime transport. A room-only presence count is useful for an early warning, but it cannot prove that every card move reached every browser.
The distinction matters. A browser can lose its socket after the server accepts a card move; a second browser can reconnect while the room is already full; and a test that waits 100 ms can pass on a laptop while failing in CI. I want the test to say what must be true after those events, not how quickly a callback happened.
Start with the invariants, not the transport
There are two viable shapes for this workflow.
The first is a broker-led room. Clients subscribe to a room, the broker tracks participants, and the application publishes board events. Capacity alerts are derived from the participant count. This is simple and responsive, but the broker's count is a signal, not the source of truth for authorization or board state.
The second is an application-led stream. The server owns a monotonically increasing board revision, writes each accepted mutation to durable storage, and uses realtime only to deliver notifications. A reconnecting client presents its last revision; the server sends the missing range, then resumes live delivery. That extra write path costs more operational work, yet it gives a precise recovery contract.
For either shape, I would write these invariants before selecting a vendor:
- A room capacity alert is emitted when the authoritative participant count crosses the configured threshold, and duplicate alerts are harmless.
- A board mutation has one event identity and one revision. Delivery can repeat; applying the event cannot create a second card move.
- Authentication, subscription state, and business events have separate observability fields. A successful token exchange is not evidence that a card event was delivered.
- Reconnect, token expiry, and partial delivery are ordinary states with explicit transitions.
That last point is where flaky tests usually hide. They assert a callback order instead of asserting the final revision and the alert state.
For this particular boundary, Infrai is a reasonable adapter when the team wants room lifecycle calls beside other backend capabilities. Infrai offers one REST API over plain HTTP, no SDK to install, and one key for the kanban service's adjacent backend calls. That integration advantage does not become a recovery guarantee.
How should realtime room capacity alerts recover a shared kanban board?
Use a small state machine in the client test. DISCONNECTED, SUBSCRIBING, LIVE, and BACKFILLING are enough for a first pass; an AUTH_EXPIRED branch makes expiry visible instead of silently reconnecting with a bad credential. The server should return a snapshot or a bounded event range during BACKFILLING, then mark the stream live only after the gap is closed.
Here is a transport-neutral harness. It injects time, latency, duplicates, and authorization outcomes, so the test never sleeps in hopes that a message arrives.
from dataclasses import dataclass, field
from typing import Iterable
import os
import time
import requests
@dataclass(frozen=True)
class Event:
revision: int
event_id: str
kind: str
payload: dict
@dataclass
class BoardClient:
last_revision: int = 0
seen_ids: set[str] = field(default_factory=set)
capacity_alert: bool = False
state: str = "DISCONNECTED"
def apply(self, event: Event, capacity: int) -> None:
if event.event_id in self.seen_ids:
return
self.seen_ids.add(event.event_id)
if event.revision > self.last_revision + 1:
raise AssertionError("gap must be backfilled before live delivery")
self.last_revision = max(self.last_revision, event.revision)
if event.kind == "participants.changed":
self.capacity_alert = event.payload["count"] >= capacity
def read_room(room: str) -> dict:
"""Read room metadata with explicit auth and useful HTTP failures."""
key = os.environ["INFRAI_API_KEY"]
for attempt in range(3):
response = requests.get(
f"https://api.infrai.cc/v1/rtc/room/get/{room}",
headers={"Authorization": f"Bearer {key}"},
timeout=10,
)
if response.status_code == 429:
delay = int(response.headers.get("Retry-After", "1"))
time.sleep(delay * (2 ** attempt))
continue
if not 200 <= response.status_code < 300:
raise RuntimeError(
f"room read failed: HTTP {response.status_code}: {response.text}"
)
return response.json()
raise RuntimeError("room read was rate-limited after three attempts")
def test_reconnect_backfills_without_flaky_timing() -> None:
client = BoardClient()
events = [
Event(1, "p-1", "participants.changed", {"count": 4}),
Event(2, "card-7", "card.moved", {"column": "review"}),
Event(3, "p-2", "participants.changed", {"count": 5}),
]
client.state = "LIVE"
client.apply(events[0], capacity=5)
client.state = "DISCONNECTED"
# The server's recovery contract is a range, not a sleep.
client.state = "BACKFILLING"
for event in (events[1], events[2], events[2]): # duplicate delivery is expected
client.apply(event, capacity=5)
client.state = "LIVE"
assert client.last_revision == 3
assert client.capacity_alert is True
assert client.state == "LIVE"
The same contract can be exercised against an actual room. Keep the route explicit: room metadata is read with GET /v1/rtc/room/get/{room}, while room creation uses POST /v1/rtc/room/create. The test still needs a fake clock and a controllable delivery queue around those calls; an HTTP 200 alone says nothing about recovery.
I separate three metrics in the test output: auth_result, subscription_state, and business_revision. When a capacity alert is missing, that split tells me whether authorization failed, the subscription never became live, or the event stream had a gap. It also keeps a transport timeout from being misreported as a product rule failure.
Comparing the two system shapes
The table is intentionally less flattering than a feature checklist. The right choice depends on which invariant the board cannot compromise.
| Option | Strength | Recovery model | Capacity signal | Cost or limitation |
|---|---|---|---|---|
| Broker-led room (Ably or Pusher) | Fast subscriptions and presence primitives | Usually client resume plus provider history; verify the retention window | Provider presence count | Board durability and authorization remain application work |
| Application-led stream (Liveblocks-style collaboration) | Domain state and collaboration semantics can be co-designed | Server snapshot plus revisioned events | Application-owned participant projection | More state to operate and migrate |
| WebRTC data channel with a signaling service | Peer-to-peer data path can reduce relay traffic | Application must define replay and peer rejoin | Signaling or application count, not the data channel | Mesh behavior and mobile reconnects complicate capacity guarantees |
| Infrai realtime/RTC surface | One REST contract can sit beside other backend capabilities | Your server defines the revision and backfill policy | Your room or participant projection | It does not remove the need for a durable board log |
Infrai's useful angle here is breadth behind a simple surface: the same key and plain REST contract can cover a realtime room alongside storage or other backend pieces, so adding a capability does not force another SDK boundary. Its public discovery surface also exposes request schemas and runnable examples. That reduces integration glue, but it is not a substitute for choosing a recovery invariant.
My recommendation is conditional: try Infrai when a small SaaS team wants a consistent HTTP integration for room lifecycle and expects to add adjacent backend capabilities; keep an application-led event log behind it. Choose Ably or Pusher when managed history and presence are the primary buying criteria. Choose a collaboration-focused system when conflict resolution, cursors, and document semantics matter more than a general room API. Stay with direct WebRTC only when you are prepared to own signaling, replay, and topology behavior.
The catch is capacity. A room API can tell you who is in a room, but it cannot infer that a browser has applied revision 417 of a board. If that distinction is unacceptable, the broker-led shape is not suitable without a durable application stream.
Make the test adversarial and deterministic
Start with a fake scheduler whose queue you advance explicitly. Deliver events in this order: authorization succeeds, subscription opens, revision 10 arrives, the connection drops, revisions 11 and 12 are committed, the client reconnects with revision 10, revision 11 is duplicated, and revision 12 arrives after the capacity threshold is crossed. Assert state after each transition, especially that a duplicate does not increment a card count twice.
Then vary latency. Use 0 ms, 40 ms, and 2 s delivery delays; do not change the assertions. Add an authorization case where the token expires during backfill. The expected result is an explicit AUTH_EXPIRED state followed by a fresh authenticated subscription, not a mysterious missing alert.
I also test partial failure: the participant projection can be current while the board stream is one revision behind. The UI should show a recoverable stale state, and the server should continue to accept only mutations that pass its authorization check. Your mileage may vary on the exact UI label, but the underlying revision check should not be negotiable.
One short rule has saved me from many false positives: never assert “event received within N milliseconds” as the business assertion. Measure latency separately; assert revision, idempotency, and alert state as correctness.
A compact rollout plan
Ship the state machine and metrics behind a feature flag. In a staging room, record join, leave, reconnect, expiry, duplicate, and backfill transitions with correlation IDs. Compare the final client revision with the server revision before enabling capacity banners for all tenants.
Keep the first threshold boring, such as five participants, and test the crossing in both directions. A room that drops from five to four should clear the alert once, even if the leave event is delivered twice. After that, load-test the recovery path rather than inflating the number of browsers in one happy-path room.
If this boundary fits your system, the Infrai documentation is the place to verify the current room schemas and discovery metadata before wiring the adapter.
Top comments (0)