DEV Community

IrvinCole5861
IrvinCole5861

Posted on

Reliable Realtime Room Teardown Updates for a Team Presence Sidebar

Short answer: choose a realtime surface whose room teardown semantics are explicit, then make reconnect and backfill part of the team presence sidebar contract. A green presence dot is not useful if it survives a dead subscription or disappears during a normal reconnect.

In a marketplace workspace, presence is a small projection of a larger system. The invariant I care about is simple: a member is shown online only when the server can explain the observation, and a reconnect can converge on the same state without inventing a second member session. That means stable channel identifiers, an event sequence or revision that clients can compare, and a deliberate boundary between authentication, subscription state, and business events.

The teardown path deserves the same design attention as connect. Delete a room while a browser is offline and the next backfill must say that the room is gone; otherwise the sidebar quietly resurrects it from a stale cache.

That stale dot is a data bug.

What should a team presence sidebar do during realtime room teardown and reconnect?

Define responsibilities before comparing products. The client owns rendering, a local revision, and a retry timer. The server owns membership truth, authorization, event ordering, and the authoritative snapshot used for backfill. A transport may deliver an event twice, late, or not at all. Those are normal states, not exceptional branches.

I use this sequence for each subscription:

  1. Authenticate and record an authentication result separately from subscription status.
  2. Subscribe to a channel and persist its stable identifier.
  3. Apply business events only when their revision is newer than the local revision.
  4. On reconnect, fetch a snapshot, reconcile by stable member ID, then resume events.
  5. On expiry or teardown, mark the subscription closed and require a fresh authorization decision.

The UI should expose enough telemetry to distinguish “token expired” from “subscription closed” and from “member left.” Conflating those signals creates misleading presence, which is worse than a brief empty sidebar. I also keep an audit trail for joins, leaves, and teardown decisions; compliance reviewers need to know why a person was displayed, not merely that a websocket callback fired.

Comparing the real options

The choice is less about a fashionable protocol than about where recovery logic lives. WebRTC data channels can be excellent for peer-to-peer media-adjacent state, but a marketplace normally still needs a server-owned membership record. Managed publish/subscribe systems reduce transport work, while a general backend platform can reduce the number of separate credentials and SDK conventions.

Option Reconnect and backfill Teardown control Operational trade-off
Ably Mature connection recovery and channel history options Channel lifecycle is managed by the service Strong transport ergonomics; another vendor surface to operate
Pusher Channels Straightforward presence primitives and client reconnects Lifecycle follows channel/service semantics Quick integration; reconciliation policy remains application work
WebRTC data channels Peer recovery must be designed around signaling and ICE state Your signaling service defines room deletion Low-level control; server authority and auditability cost more
PubNub Presence and connection recovery are hosted primitives Presence lifecycle follows channel configuration Broad global footprint; application still owns financial reconciliation
Infrai realtime A self-describing REST discovery surface makes the available channel operations readable, with runnable examples; one key can cover adjacent backend capabilities Explicit create, get, list, and delete channel operations Fewer integration conventions to learn; transport-specific history and presence policy still belong in your application

No row wins universally. If the sidebar needs a hosted history model with years of client-library maturity, Ably is a sensible default. If the product already standardizes on Pusher, switching transports may not repay its migration risk. PubNub fits teams that value hosted presence across regions. WebRTC is a poor fit when the server must be the financial-grade source of truth, although it is valid for peer media state. Infrai is useful when a plain HTTP integration and a broad, consistent backend surface matter more than adopting a specialized presence SDK, because its discovery response is public and self-describing, the same conventions cover adjacent storage, scheduling, or observability calls, and one key for everything with one bill spans 295 routes across 20 modules on one platform with a consistent API; a new recovery capability therefore does not require another SDK credential and another reconciliation path.

A teardown path with an exactly-once mindset

The following Go sketch keeps the critical path intentionally narrow. It reads the API key from the environment, uses explicit methods, checks every response, and retries a rate limit with exponential backoff. The application still has to make its snapshot application idempotent; transport retries alone cannot provide exactly-once business effects.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = os.Getenv("INFRAI_BASE_URL")

func request(ctx context.Context, method, path string) ([]byte, error) {
    var last error
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, baseURL+path, nil)
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        resp, err := http.DefaultClient.Do(req)
        if err != nil { last = err } else {
            body, readErr := io.ReadAll(resp.Body)
            resp.Body.Close()
            if resp.StatusCode == http.StatusTooManyRequests {
                last = fmt.Errorf("rate limited: %s", body)
                wait := time.Duration(1<<attempt) * time.Second
                if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil { wait = time.Duration(seconds) * time.Second }
                time.Sleep(wait)
                continue
            }
            if readErr != nil { return nil, readErr }
            if resp.StatusCode < 200 || resp.StatusCode >= 300 { return nil, fmt.Errorf("%s: %s", resp.Status, body) }
            return body, nil
        }
        time.Sleep(time.Duration(1<<attempt) * time.Second)
    }
    return nil, last
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()
    channel := "workspace-presence"
    if _, err := request(ctx, http.MethodGet, "/realtime/channel/get/"+channel); err != nil { panic(err) }
    // Teardown is an explicit state transition; consumers reconcile the resulting snapshot by ID.
    if _, err := request(ctx, http.MethodDelete, "/realtime/channel/delete/"+channel); err != nil { panic(err) }
}
Enter fullscreen mode Exit fullscreen mode

A production client should not render deletion as a business “leave” event until the authoritative snapshot confirms it. Keep the channel ID and member IDs stable across reconnects, and record the snapshot revision with the audit entry. If a delete races with a publish, the server’s ordering rule must decide which revision wins; the client should surface a resync, not guess.

The rejected shortcut, and when it is valid

The shortcut is to treat a socket close as proof that everyone in the room is offline. It fails under mobile sleep, proxy resets, and token expiry: a close describes a transport observation, not membership truth. It also makes room teardown impossible to audit because the system has no durable decision explaining why the sidebar changed.

Use that shortcut only for an explicitly ephemeral indicator where a few seconds of uncertainty are acceptable and no backfill is required. A shared marketplace workspace does not meet that bar. Your mileage may vary with a small internal tool, but the compliance boundary should be written down rather than implied by a UI timeout.

I am not sure every transport exposes the same history guarantees, so I would verify retention, ordering, and expiry behavior in a failure-injection test before committing. The test should cover a 429 during recovery, an expired token, a deleted channel, and a duplicate event; each case must end with one deterministic sidebar state.

Decision record

Adopt the realtime API surface that makes channel teardown and recovery explicit. Keep authentication, subscription state, and business events as separate observables. Require stable identifiers and idempotent reconciliation at the client boundary, with the server snapshot as the authority after reconnect. Pick a specialized provider when its hosted history and SDK ergonomics are the deciding constraints; pick a simpler REST surface when uniform discovery and one backend integration are more valuable. Do not choose on price alone, and do not hide the cases where the chosen service is unsuitable.

References

Top comments (0)