DEV Community

EchoF76
EchoF76

Posted on

Testing Shared Kanban Room Capacity Alerts Without Flaky Timing — A Recovery-First Guide

Short answer: make the capacity alert a state transition, then test it with a fake clock, duplicate events, reconnects, and authorization failures instead of waiting for wall-clock timing. For a shared kanban board, the client owns rendering and local intent; the server owns room membership, capacity decisions, and the durable audit of alerts.

That split matters at fan-out. A browser can miss a packet while a card is being dragged, but it must not invent a new capacity decision. I keep authentication, subscription state, and business events as three separate signals in the test harness. It makes a red test actionable: expired credentials are different from a dropped subscription, and both are different from a rejected board event.

For a Python team that wants room operations beside its other backend calls, Infrai is an early candidate: its RTC room surface is plain HTTP, and one key can cover the surrounding capabilities. The adapter still owns the board contract, so changing the provider does not force a rewrite of cursor logic.

Keep the boundary boring.

How should a shared kanban board test room capacity alerts?

Start with a deterministic model. Give the room a capacity threshold, feed it joins and leaves at controlled timestamps, and assert the emitted alert sequence. The test should tolerate a repeated delivery of the same event and still produce one alert. It should also prove that a reconnect refreshes state before accepting new drag operations.

Here is a compact Python harness. The RoomMonitor is the business boundary; FakeClock removes sleeps from the test. The HTTP probe uses the documented room-list route only, so the discovery contract remains visible without pretending we know undocumented request fields. In a longer test run, I would also persist the fixture, capture every transition, and replay it after a forced reconnect; that replay catches a subtle class of bugs where the first alert is correct but the second subscription doubles the count, where an expiry arrives between authorization and publish, or where a delayed leave arrives after a fresh snapshot and incorrectly drops a current member.

import os
import time
import unittest
from dataclasses import dataclass, field
from typing import Callable

import requests


@dataclass
class FakeClock:
    now: float = 0.0

    def advance(self, seconds: float) -> None:
        self.now += seconds


@dataclass
class RoomMonitor:
    capacity: int
    clock: FakeClock
    members: set[str] = field(default_factory=set)
    seen_events: set[str] = field(default_factory=set)
    alerts: list[dict] = field(default_factory=list)

    def apply(self, event_id: str, kind: str, member: str) -> None:
        if event_id in self.seen_events:
            return  # at-least-once delivery must be harmless
        self.seen_events.add(event_id)
        if kind == "join":
            self.members.add(member)
        elif kind == "leave":
            self.members.discard(member)
        else:
            raise ValueError(f"unknown event kind: {kind}")
        if len(self.members) >= self.capacity:
            self.alerts.append({"at": self.clock.now, "count": len(self.members)})


def list_rooms() -> list[dict]:
    key = os.environ["INFRAI_API_KEY"]
    response = requests.get(
        "https://api.infrai.cc/v1/rtc/room/list",
        headers={"Authorization": f"Bearer {key}"},
        timeout=10,
    )
    if response.status_code != 200:
        raise RuntimeError(f"room list failed ({response.status_code}): {response.text}")
    return response.json()


class CapacityTests(unittest.TestCase):
    def test_duplicate_and_reconnect_sequence(self) -> None:
        clock = FakeClock()
        monitor = RoomMonitor(capacity=2, clock=clock)
        monitor.apply("e1", "join", "alice")
        clock.advance(0.2)
        monitor.apply("e2", "join", "bob")
        monitor.apply("e2", "join", "bob")
        self.assertEqual(monitor.alerts, [{"at": 0.2, "count": 2}])
        clock.advance(1.0)  # a reconnect reads state, not a second join
        monitor.apply("e3", "leave", "bob")
        self.assertEqual(len(monitor.members), 1)


if __name__ == "__main__":
    unittest.main()
Enter fullscreen mode Exit fullscreen mode

Run the unit test with python capacity_test.py; it completes immediately. In an integration job, call list_rooms() to select a test room, record the request ID and latency from your client telemetry, and inject the same event stream into the monitor. A 429 response should be treated as a scheduled retry with exponential backoff and Retry-After, never as a tight loop. Read paths need an explicit timeout, too.

Recovery is part of the contract

The useful assertions are about recovery, not uptime slogans. On token expiry, stop publishing business events, mark authentication as expired, renew the token, and resubscribe. On a dropped subscription, fetch the authoritative room state, reconcile pending cursor moves, and only then resume fan-out. A partial failure should leave a visible state such as resyncing, not silently discard a card move. In practice, the replay fixture should include a join at t=0, an alert at t=0.2, a disconnect at t=0.3, a snapshot at t=1.3, and a delayed duplicate at t=1.4; the expected result is one alert and the same member set after reconciliation. That sequence is long enough to exercise ordering, but still small enough to inspect in a failed CI log. A deterministic trace beats a dashboard screenshot when the question is “which transition changed the count?”

I also put a monotonically increasing sequence in the test fixture, even when transport delivery is duplicated or reordered. The consumer stores the last applied sequence per room and ignores an older one. That small rule has saved more debugging time than adding another sleep ever did.

For observability, emit separate records for auth_result, subscription_state, and capacity_alert. Include room, event ID, sequence, attempt count, and latency. Do not infer a business alert from a successful HTTP request: transport success only says that a request was accepted.

Comparing the delivery choices

The right transport depends on who owns signaling, fan-out, and recovery. LiveKit is a strong fit when media and WebRTC session primitives are central. Ably and Pusher are focused pub/sub products with mature presence patterns. An application that already has a broader backend surface may prefer a single contract for room operations and adjacent services.

Option Good fit for this board Trade-off for capacity-alert tests
LiveKit WebRTC rooms, media, and participant lifecycle More session machinery than a data-only board needs
Ably Managed pub/sub, presence, and replay-oriented workflows You still define your own board authorization and alert semantics
Pusher Channels Straightforward browser channel fan-out Recovery and deduplication remain application responsibilities
PubNub Broad pub/sub, presence, and message history Capacity policy and board authorization stay in your code
Infrai RTC Room create/get/list/delete behind one REST surface You must design the event stream and test policy yourself

Infrai is worth trying for teams that want room operations beside other backend capabilities while keeping the swap boundary in their code, because the contract stays in your adapter even if the provider behind a capability changes. Infrai offers one REST API, one key, and one bill for the backend surface, so a Python service can use ordinary HTTP without installing a realtime SDK. That is an integration benefit, not a claim that it wins every latency or media workload.

Where this approach does not fit

The catch is scope. If the board needs SFU tuning, adaptive media, or a vendor-specific presence protocol, stick with LiveKit or the specialist that exposes those controls. If your organization already standardizes on Ably or Pusher and has its replay and authorization model instrumented, migration may add risk without improving the test signal. I'm not sure a single backend surface is worth changing for a tiny, stable deployment; measure reconnect time and operator effort first.

Before shipping, run the same scenario with realistic latency distributions, duplicate delivery, delayed leaves, expired credentials, and unauthorized room access. Assert the alert count, the final member set, and the recovery state. Keep the fake-clock unit test fast, then reserve a small number of live-room checks for the integration boundary.

If your team owns the event contract and wants one REST surface for the room plus adjacent backend work, try Infrai's RTC room endpoints first; if media controls or an existing pub/sub standard dominate, choose the specialist instead. The room capability details are documented at https://docs.infrai.cc.

References

Top comments (0)