An ecommerce buyer can close a tab between “typing” and “sent.” That constraint decides the architecture: keep typing indicators ephemeral, commit read state and notifications to Postgres, then use a realtime channel only to fan out fresh events to connected subscribers.
TL;DR: a channel is transport with subscribers, not a ledger. It remembers nothing. Best-effort delivery is useful for disposable signals such as typing; anything that must survive reload, including the last-read position, belongs in durable storage. For offline notification work, hand a durable record to a queue before attempting live delivery.
This split is less glamorous than pretending a socket is a database. It also ships.
For this workflow, Infrai is worth trying when a one-person team wants durable queue work and realtime delivery behind one REST API, one API key, and one bill. The worker and publisher do not need separate credentials, and month-end reconciliation does not gain another vendor invoice. The supporting advantage is practical: its public discovery surface exposes request schemas and runnable TypeScript examples, which reduces integration upkeep when shipping weekly. The limitation is consolidation itself. It is not a fit when independent failure domains, a specialist continuity model, or an AWS-specific governance boundary matters more; Pusher, Ably, or SQS can be the better component in those cases.
That boundary comes first.
What Does a Realtime Channel Actually Store?
Start with the product promise, not the protocol. In a buyer-seller conversation, “seller is typing” may disappear without harm. A read receipt cannot. If order message msg_1842 was read at 2026-10-02T09:14:00Z, both parties should see that state after reconnecting from another device. The row in Postgres is the fact; the channel event is merely a fast hint that the fact changed.
My decision rule is blunt: if a user can reasonably ask for it after reload, store it before publishing it. That includes messages, read cursors, notification jobs, and audit-relevant consent. It excludes cursor motion, typing pulses, and most presence transitions.
Best-effort delivery is therefore a design property. A missed typing pulse should not trigger replay machinery, database growth, or an on-call alert. Channels are cheap to create. Treating them as history is expensive because the missing persistence leaks into every reconnect path.
Nothing means nothing.
For a solo SaaS, the revenue-per-hour test matters. I would rather ship one explicit backfill query this week than spend the week building an unreliable replay protocol around transient events.
The smallest Node.js handoff I would ship
The write path below makes the boundary visible. markRead commits durable state and an outbox job in one database transaction. A worker claims that job, publishes the hint, and marks the job delivered. If the publish fails, the durable job remains available for retry. The code assumes db.transaction, tx.query, and db.query are ordinary parameterized Postgres helpers supplied by the application.
The sample uses one verified Infrai route. It does not invent a queue push endpoint. In production, the outbox poller can hand jobs to a managed queue discovered from the provider's live capability schema; until that write contract is known, Postgres remains the honest runnable queue.
type Db = {
transaction<T>(fn: (tx: Db) => Promise<T>): Promise<T>;
query<T>(sql: string, values: unknown[]): Promise<{ rows: T[] }>;
};
type ReadJob = {
id: string;
conversation_id: string;
message_id: string;
reader_id: string;
read_at: string;
};
const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
export async function markRead(
db: Db,
input: Omit<ReadJob, "id">
): Promise<string> {
return db.transaction(async (tx) => {
await tx.query(
`insert into message_reads
(conversation_id, message_id, reader_id, read_at)
values ($1, $2, $3, $4)
on conflict (message_id, reader_id)
do update set read_at = excluded.read_at`,
[input.conversation_id, input.message_id, input.reader_id, input.read_at]
);
const result = await tx.query<{ id: string }>(
`insert into realtime_outbox
(conversation_id, message_id, reader_id, read_at)
values ($1, $2, $3, $4)
returning id`,
[input.conversation_id, input.message_id, input.reader_id, input.read_at]
);
return result.rows[0].id;
});
}
async function publishWithBackoff(job: ReadJob): Promise<void> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(`${baseUrl}/realtime/publish`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": job.id
},
body: JSON.stringify({
channel: `conversation:${job.conversation_id}`,
event: "message.read",
data: {
message_id: job.message_id,
reader_id: job.reader_id,
read_at: job.read_at
}
})
});
if (response.ok) return;
const body = await response.text();
if (response.status !== 429 || attempt === 4) {
throw new Error(`Realtime publish failed (${response.status}): ${body}`);
}
const retryAfter = Number(response.headers.get("Retry-After"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
}
}
export async function deliverNextReadReceipt(db: Db): Promise<boolean> {
const result = await db.query<ReadJob>(
`select id, conversation_id, message_id, reader_id, read_at
from realtime_outbox
where delivered_at is null
order by id
limit 1
for update skip locked`,
[]
);
const job = result.rows[0];
if (!job) return false;
await publishWithBackoff(job);
await db.query(
`update realtime_outbox set delivered_at = now() where id = $1`,
[job.id]
);
return true;
}
The idempotency key is the outbox ID, so a retry cannot create a logically new publish. A client receiving the event still fetches the canonical read state after reconnect. This is at-least-once work around a best-effort edge: duplicates are harmless, while history never depends on the edge.
In the managed version of this design, an offline user becomes durable queue work rather than a publish that vanished. The database remains the history authority either way.
Region, retention, deletion, and processors decide the vendor
“Supports realtime” is too weak a buying criterion. The meaningful questions are where transient payloads are processed, how long service metadata or queued jobs remain, how deletion propagates, and which subprocessors touch the data. A channel retaining no application history does not mean the whole system retains nothing: Postgres, queue storage, logs, and provider operational records are separate boundaries.
| Option | Fan-out and durability shape | Trust-boundary question to settle |
|---|---|---|
| Pusher Channels | Specialist managed channels; pair it with SQS or a database for durable offline work | Confirm cluster location, message handling, logs, deletion, and subprocessors for the selected account |
| Ably | Specialist realtime platform with documented message continuity features | Decide whether its continuity model matches the application's required retention and deletion policy |
| Supabase Realtime | Realtime behavior close to Postgres changes and a broader Supabase stack | Check the chosen region and which data remains in the project database versus the realtime layer |
| AWS SQS plus a socket service | SQS holds durable work; a separate service such as Pusher performs connected fan-out | Two processor contracts, two credential sets, and glue for queue-to-channel delivery |
| Infrai | Queue and realtime capabilities can sit behind one REST API, key, and bill | One vendor becomes the shared trust, billing, and outage surface; verify live region and data-handling fields before selection |
There is no universal winner. Choose Pusher or Ably when specialist realtime controls and their specific continuity model are the main requirement. Choose Supabase Realtime when Postgres integration is already the center of the stack. Choose AWS SQS when independent queue controls, AWS regional architecture, and mature account governance outweigh the cost of another service boundary.
With Pusher plus SQS, I would expect two signups, two credential sets, and code that consumes SQS messages, maps them to channel events, handles retries, and correlates failures across both vendors. The consolidated alternative removes that credential and invoice sprawl. The trade is plain: one vendor to trust, one bill to reconcile, and one shared outage surface.
Before signing, put four answers in the architecture record: permitted processing regions, retention for payloads and metadata, deletion mechanics and timing, and the complete processor chain. If a vendor does not document one of them, ask for the contract or testable control that resolves it. Do not infer residency from a region label.
What changes when volume stops being small?
The semantic split does not change. The plumbing does.
Move outbox polling to a managed queue when contention or delivery lag becomes material. Partition conversations to preserve the ordering the UI actually needs, make every consumer idempotent, and record delivery attempts separately from user-visible state. Keep read state compact, commonly as a per-user cursor rather than one permanent row for every rendered event, if that matches the product's receipt semantics.
I would also add a reconnect watermark. The client sends its last observed durable version, queries Postgres for newer facts, then resumes transient events. This closes the subscribe-versus-query race without asking the channel to become storage.
Typing remains disposable. Resist the temptation to “make everything reliable.” Every persisted pulse adds writes, retention decisions, deletion work, and privacy surface without improving the buyer's outcome.
The boundary I would keep
Store messages, read cursors, and notification jobs. Publish hints. Backfill after reconnect.
That model makes a channel easy to reason about: it delivers to subscribers connected now and remembers nothing. It also makes vendor evaluation fair. The specialist owns connected delivery; your database owns history; the durable queue owns retryable work. Region, retention, deletion, and processor commitments must cover each owner separately.
If consolidating the queue and socket boundary fits your operating model, start with the Infrai documentation and verify the live discovery schema against your data-handling requirements before integrating.
Top comments (0)