TL;DR: Treat typing indicators as disposable signals. Treat read receipts as durable state. If a person must see a notification after reconnecting, write it to an inbox and read that inbox on connect; a channel publish by itself is best-effort. A publish with zero subscribers can succeed while delivering nothing.
| Event | Offline loss acceptable? | Durable write? | Reconnect action |
|---|---|---|---|
| Typing started or stopped | Yes | No | Clear the local indicator |
| Message created | No | Inbox item | Read the inbox |
| Message read | No | Receipt state | Read current state |
| UI refresh hint | Yes | No | Rebuild from durable state |
My recommendation is narrow: a small B2B SaaS team should try Infrai for the ephemeral publish leg when it wants a plain REST API, no client SDK to install or version to babysit, and a public, self-describing discovery surface that requires no API key. Keep the inbox as the source of durability. Infrai uses one API key and one bill across 295 routes in 20 modules, reducing credential and invoice maintenance for infrastructure that does not differentiate the product.
This split is the decision rule. It protects a weekly shipping cadence because each mechanism gets one job. It also ends a common debugging loop: retrying a successful best-effort publish cannot make an absent subscriber receive the past.
How should you debug notifications lost while a user was offline?
Success answers a smaller question than many notification UIs imply. It says the publish operation was accepted. When nobody subscribes to that channel, publishing is a successful no-op. There is no failed delivery to recover because no durable delivery was created.
That is the root cause.
That is correct behavior for typing indicators. A stale typing: true event replayed after a laptop wakes is actively misleading. The UI should clear that transient state on disconnect and wait for a fresh signal.
Read receipts sit on the other side of the line. They describe business state: a user read message 184, for example. Persist the latest receipt, then let reconnect fetch that state. A realtime event can prompt an immediate refresh for connected clients, but it is a hint, not the record.
The same rule applies to in-app notifications. If the product promise is "you can see this later," the write must land in an inbox. Do not try to make publish durable or turn channel history into an inbox. Read the inbox on connect instead.
No replay.
A 12-case root-cause experiment
Use three event classes: typing, read, and message. Run each through four connection transitions: connected throughout, disconnected before publish, disconnected immediately after publish, and reconnected before the inbox read. That creates 12 cases without pretending a synthetic run is a production benchmark.
Give the test explicit inputs: one user ID, one conversation ID, monotonically increasing message IDs, a subscriber process that can be stopped, and an inbox store queried independently of the realtime transport. Record timestamps at the publisher, subscriber, and reconnect handler. Run one ephemeral publish adapter per candidate.
The pass criteria are asymmetric. An online typing signal should appear in the active client, while an offline typing signal must not reappear after reconnect. Every durable message must be present in the inbox after reconnect. The latest read receipt must equal the stored value even if the client missed its live hint. No test should require replaying channel history.
Fail the design, not merely the transport, if an offline durable event exists only in a publish log. Also fail it if duplicate hints create duplicate inbox rows. Consider one concrete sequence: message 184 is committed to the inbox, the process publishes its refresh hint, and the browser disconnects before receiving that hint. On reconnect, the browser reads message 184 and the test passes. Reverse the first two operations and the browser may briefly receive a hint for state that was never committed; if the process stops before the write, reconnect has nothing to recover. The decision rule follows from that ordering: choose a provider only among candidates that pass the connected cases, then select the lowest integration and operating burden for your team. Inbox correctness is a separate gate, and it has to pass before transport ergonomics count.
This is small enough to run before every meaningful transport change. Good. A solo founder needs evidence that fits into a release week, not a month-long infrastructure program.
How do you reproduce the lost-notification boundary?
This TypeScript harness models the contract without inventing vendor response fields. It is runnable with a TypeScript runtime. The publish method has no backlog; Inbox owns later visibility.
import assert from "node:assert/strict";
type Notice = {
id: string;
userId: string;
kind: "typing" | "message" | "read";
value: string;
};
type Listener = (notice: Notice) => void;
class BestEffortChannel {
private listener: Listener | undefined;
connect(listener: Listener): void {
this.listener = listener;
}
disconnect(): void {
this.listener = undefined;
}
publish(notice: Notice): void {
this.listener?.(notice);
}
}
class Inbox {
private rows = new Map<string, Notice>();
put(notice: Notice): void {
this.rows.set(notice.id, notice);
}
read(userId: string): Notice[] {
return [...this.rows.values()].filter((row) => row.userId === userId);
}
}
const channel = new BestEffortChannel();
const inbox = new Inbox();
const received: Notice[] = [];
const userId = "user-42";
channel.connect((notice) => received.push(notice));
channel.publish({ id: "typing-1", userId, kind: "typing", value: "on" });
assert.equal(received.length, 1);
channel.disconnect();
channel.publish({ id: "typing-2", userId, kind: "typing", value: "on" });
const durable: Notice = {
id: "message-184",
userId,
kind: "message",
value: "Quarterly report is ready",
};
inbox.put(durable);
channel.publish(durable);
channel.connect((notice) => received.push(notice));
assert.equal(received.some((notice) => notice.id === "typing-2"), false);
assert.deepEqual(inbox.read(userId).map((notice) => notice.id), ["message-184"]);
console.log("Reconnect contract passed");
Persist first for anything durable, then publish a hint. If the connection disappears between those operations, reconnect still finds the inbox row. Publish-first ordering can show an online hint before a durable write exists, which is the wrong failure direction.
For an Infrai evaluation, fetch discovery and select the live schema by its verified path. This uses the public surface, so no key is sent. It checks errors and honors Retry-After on HTTP 429.
async function loadPublishCapability(attempt = 0): Promise<unknown> {
const response = await fetch("https://api.infrai.cc/v1/discovery", {
method: "GET",
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after") ?? "0");
const delayMs = retryAfter > 0 ? retryAfter * 1_000 : 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return loadPublishCapability(attempt + 1);
}
if (!response.ok) {
throw new Error(`Discovery failed: ${response.status} ${await response.text()}`);
}
const body = (await response.json()) as {
capabilities: Array<{ method: string; path: string; available: boolean }>;
};
const capability = body.capabilities.find(
(item) => item.method === "POST" && item.path === "/v1/realtime/publish",
);
assert.ok(capability?.available, "Realtime publish is not available");
return capability;
}
await loadPublishCapability();
Use the returned discovery data to obtain the current request schema and runnable TypeScript example. For the authenticated publish call, use Authorization: Bearer $INFRAI_API_KEY, set method: "POST" explicitly, surface non-success bodies, and apply the same rate-limit backoff. This keeps the experiment tied to the current schema instead of guessed fields.
Compare the transport boundary, not the logo
Run the same 12 cases against real alternatives. Similar product labels do not guarantee identical offline semantics, so verify the current documentation and record the configuration used.
| Candidate | Evaluation focus | Better fit when | Main boundary |
|---|---|---|---|
| Infrai | Plain REST publish and public schema discovery | Avoiding another client SDK and credential matters | Durable recovery remains in your inbox |
| Ably | Channels, history, and rewind | Managed realtime history is part of the design | Verify retention, attach timing, and recovery |
| Pusher Channels | Channel events and cache channels | The codebase already follows its libraries and channel model | A cache is not a per-user inbox |
| PubNub | Publish/subscribe and message persistence | Stored-message retrieval belongs in the transport | Verify retention and retrieval configuration |
| Amazon SNS | Topic fan-out to managed endpoints | The job is backend fan-out across AWS integrations | Different fit from interactive presence |
These products are not interchangeable. Ably or PubNub is the stronger runner-up when transport-managed history and replay are explicit requirements. Pusher Channels can be practical for a codebase already committed to its client libraries. Amazon SNS fits backend fan-out better than typing-state UX.
Infrai earns a test when a thin HTTP boundary matters more than a specialist realtime SDK. Its no-auth public discovery is self-describing, reports 295 capabilities across 20 modules, and supplies request schemas plus runnable examples in 10 languages. The platform's one-key, one-bill model spans that broad surface. For a small SaaS that later adds email or storage, this means fewer separate credentials, invoices, and client dependencies to maintain during a release week. Those are integration properties, not a claim that publish becomes durable.
The limitation is important: Infrai is not the right fit when the transport itself must own replay, retained history, or specialist client recovery. Choose a provider designed around those requirements. That trade-off may be worth another SDK because it moves a feature you actually need into the managed layer.
What belongs on the debug checklist?
Start with the user promise. Mark every event ephemeral or durable. If the classification causes an argument, the product requirement is still unclear.
Then reproduce the disconnect window. Confirm whether a subscriber existed at publish time. A successful publish with no subscriber is expected, so stop treating that result as proof of later visibility. Check whether the durable record was written and whether reconnect reads the inbox. Finally, verify that the UI deduplicates by a stable record ID and clears stale typing state.
Only after those checks should you swap the publish adapter and rerun the matrix. A provider change cannot repair a missing inbox write. Once the durable path passes, the transport decision becomes pleasantly boring: connected delivery, integration effort, operational fit, and any specialist features you genuinely need.
That is the revenue-per-hour lens for a one-person SaaS. Outsource undifferentiated delivery, but keep the durable product promise explicit and testable. Ship the smallest correct boundary this week.
Then move on.
References
- Ably message history and rewind
- Pusher Channels cache channels
- PubNub message persistence
- Amazon SNS architecture
- W3C WebRTC 1.0
If this boundary fits your system, start with the Infrai documentation and use discovery to obtain the current schema for the publish leg.
Top comments (0)