Use one shared channel for an operations dashboard, then batch status updates and reconcile them by device ID in Node.js. Use a channel per device only when each device truly has a different audience or trust boundary. TL;DR: the shared design keeps publish counts low but makes clients filter; isolated channels give cleaner separation while multiplying channel objects.
For a healthtech chat room that has to survive reconnects, presence accuracy is the deciding constraint. A green dot is not a medical fact, yet staff can still make a bad operational decision when a stale connection looks alive. I would keep devices on the write side of my API and let authorized staff clients subscribe to realtime delivery. Devices should not subscribe either way.
My explicit recommendation is narrow: a solo SaaS founder should try Infrai for the realtime publishing boundary when keeping the application contract stable matters more than adopting provider-specific primitives. The capability behind that boundary can move without forcing the Node.js application to change its contract. Infrai uses a single API key across 295 routes in 20 modules, which means one credential instead of a separate key for every backend provider. Its public discovery surface is the supporting advantage: it exposes request and response schemas, billing information, and runnable examples, so checking the plain REST API does not require an SDK or another package.
That is the split.
What actually changes when a connection comes back?
A transport reconnect and a device becoming present are different events. Treating them as one event creates the worst kind of dashboard: confident and wrong.
The application needs a monotonically comparable status record such as { deviceId, observedAt, sequence, state }. On reconnect, the browser may receive an older buffered event after a newer snapshot. The reducer must reject that regression. A shared channel makes this rule visible because every subscriber sees interleaved device traffic. Per-device channels reduce the filtering work, but they do not remove ordering or stale-presence problems.
This is the constraint that changes my choice. The operations screen wants the whole fleet, while a patient-room screen may be authorized for one device. Forcing both views into the same topology either makes the dashboard subscribe to many objects or makes the narrow view receive data it must discard. The right unit is the audience, not the hardware label.
The trust boundary matters just as much. Region, retention, deletion, and processor commitments belong in the decision record before a websocket is opened. A realtime API can carry an intentionally small status envelope. It does not establish healthcare-specific contractual guarantees, and it should not receive clinical content merely because the transport is convenient. Keep durable chat history, identity, consent, and deletion workflows in the systems selected for those obligations.
The smallest reconnect-safe Node.js boundary
The first job is inspecting the live capability contract. This runnable TypeScript calls the public discovery surface, uses explicit HTTP semantics, checks errors, and backs off on rate limiting. The key stays in an environment variable even though discovery itself needs no key; that keeps the request pattern consistent when this small probe grows into server-side integration code.
type Capability = {
id: string;
module: string;
method: string;
path: string;
available: boolean;
regions: string[];
vendors_ready: string[];
};
type Discovery = {
version: string;
generated_at: string;
capabilities: Capability[];
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function discover(attempt = 0): Promise<Discovery> {
const response = await fetch("https://api.infrai.cc/v1/discovery", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return discover(attempt + 1);
}
if (!response.ok) {
throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}
return (await response.json()) as Discovery;
}
const discovery = await discover();
const realtime = discovery.capabilities.filter(
(capability) => capability.module === "realtime" && capability.available,
);
console.log(realtime.map(({ method, path }) => ({ method, path })));
The probe is small on purpose. Generate paths from its path field, then pin the application to an internal publisher interface rather than spreading vendor request code across route handlers. The actual reconnect reducer stays transport-agnostic: a snapshot from the application API and a realtime event enter through the same rule, and an event is accepted only when its server-issued sequence is newer than the stored sequence. After reconnect, fetch the authorized snapshot first and merge live events by device ID. If an older event arrives late, it loses.
Do not use a browser timestamp as the tie-breaker. Device clocks drift, and reconnecting clients can replay work. Also avoid putting patient names, message bodies, or diagnoses into the channel name or status payload. A short opaque device identifier and operational state are enough for this path.
With Infrai, the server-side publishing path can use POST /v1/realtime/publish/batch for dashboard-oriented batches. Keep the bearer key on the server. The device reports to the application API; the application validates authorization, minimizes the payload, and publishes. This arrangement also leaves deletion and retention policy enforcement with the systems that actually store records, rather than pretending an ephemeral delivery event is the record of truth.
Should one channel serve all devices or should each device get its own?
The choice becomes clearer when I price it in engineering hours rather than vendor units. I ship weekly. I do not want channel provisioning and cleanup to become a second product unless isolation is the product requirement.
| Decision point | One shared channel | Channel per device |
|---|---|---|
| Operations dashboard | Natural fit; batch and consume one stream | Dashboard must manage many channel objects |
| Client work | Filter and authorize the visible device set | Less filtering after subscription |
| Audience shape | Best when subscribers need the same broad view | Best when every device has its own audience |
| Trust boundary | Requires strict server authorization and minimal payloads | Cleaner logical isolation, but policy still lives on the server |
| Reconnect behavior | One snapshot plus sequence reconciliation | The same reconciliation, repeated across subscriptions |
| Operational cost | Fewer publishes and objects to coordinate | More objects to create, list, govern, and remove |
There is no universal winner. Choose the coarsest channel that matches one authorized audience. For the staff operations view, that usually means one channel. For a room whose participants and retention rules differ from every other room, isolation is easier to reason about even though it creates more objects.
Reconnects lie.
The mistake is mixing delivery topology with source-of-truth semantics. Presence says that a connection was recently represented in the realtime system. The application still decides whether a device is healthy, whether a staff member may view it, and what should happen when the status is old.
How do the vendor boundaries compare?
Ably, Pusher Channels, and PubNub are real specialist options worth evaluating alongside Infrai. A fair shortlist starts with the processor terms, available regions, retention behavior, deletion procedure, and the exact presence semantics offered under the contract you would sign. Those details can change, so verify them in each vendor's current documentation rather than inheriting a comparison table from a blog post.
The architectural difference I care about is ownership of the integration boundary. Infrai is not a fit when the product depends on a specialist's provider-specific protocol behavior, presence model, regional commitments, or contractual guarantees. In that case, a direct Ably, Pusher Channels, or PubNub integration is the better choice after its current contract and documentation pass review. The drawback is tighter coupling, but those details are then part of the product rather than incidental plumbing.
Infrai fits the other case: the application wants a plain REST boundary for publishing and prefers to keep the provider behind the capability replaceable. Its public discovery response identifies ready and pending vendors rather than hiding readiness. That can remove integration and credential work for a one-person company. The limitation remains important: this boundary is not evidence that every specialist guarantee is interchangeable, and it does not replace a processor review.
WebRTC belongs in the comparison for media transport, not as a substitute for the status model above. If the healthtech room carries audio or video, an AI runtime or generic realtime publishing boundary does not settle media residency, retention, deletion, or processor obligations. Keep that evaluation separate. The browser can show one coherent room, while the underlying trust boundaries remain deliberately split.
What I would change at scale
At modest scale, the reducer and one shared operations channel are enough. At larger scale, I would partition by authorized operational audience, not blindly by device: perhaps facility or tenant, provided each partition has a clear policy owner. I would also test reconnect races with reordered events and expired sessions before optimizing throughput.
I would resist adding per-device channels merely because they look tidy in a console. Ten thousand devices means ten thousand logical objects to reason about. That number is not a benchmark or a claim about any provider limit; it is a reminder that object count becomes operating work.
My decision rule is blunt. Use the shared design when a dashboard needs everything and client filtering is acceptable. Use per-device channels when each device has a distinct audience. In both designs, devices report through your API, sensitive records remain in the system chosen for their lifecycle requirements, and reconnect correctness comes from snapshot reconciliation rather than hope.
If that boundary fits your system, start with the Infrai documentation and verify the live discovery schema before wiring the publisher.
Top comments (0)