DEV Community

FrostY45
FrostY45

Posted on

Node.js Queue Recovery for Stuck Seller Notifications (Per-Recipient Status Polling)

Render each seller's new-order message before a batch of email and SMS event notifications enters the delivery queue, then track one state machine per recipient and channel. The deciding constraint is template ownership: if the application owns rendering, a retry can reuse an immutable payload; if a downstream provider owns it, a later retry may render different content from the same event.

TL;DR: Do not poll or retry a batch as one unit. Give every (order, seller, channel) delivery a stable idempotency key, record accepted and terminal outcomes separately, and retry only recipients whose last classified result is retryable. A batch is an observation boundary, not a correctness boundary.

That rule prevents the ugliest recovery pattern: 98 sellers receive the alert, two addresses fail temporarily, and a whole-batch replay sends 98 duplicates. It also makes a queue that merely looks stuck distinguishable from one whose workers have actually stopped.

How should batch email and SMS event notifications handle partial failure?

Start with the signal. A single processing count hides at least four states: waiting for a worker, handed to a transport, accepted by that transport, and awaiting a later delivery report. Those states require different responses. Queue age can indicate worker starvation; an old accepted message can indicate delayed or missing status reconciliation. Neither observation proves that the recipient received nothing.

Email adds an important classification boundary. SMTP reply codes beginning with 4 indicate a transient negative completion, while 5 indicates a permanent negative completion. Treating both as generic failure either burns retries on permanent rejection or drops recoverable work. Acceptance is still not inbox placement: RFC 5321 defines transfer behavior, not what a mailbox UI eventually shows.

SMS transports commonly acknowledge submission before producing a later recipient-level outcome. Model that callback or poll result as a transition on the existing delivery record, never as permission to create a fresh send. Reports can arrive late or more than once. The transition must therefore be idempotent and reject regressions from a terminal state.

For this marketplace flow, the useful unit is narrow:

Field Example Operational purpose
Event order_7319 Ties notifications to the business action
Recipient seller_204 Isolates one seller from the batch
Channel email or sms Keeps channel outcomes independent
Template version new-order-v7 Makes replay content explainable
Attempt 3 Bounds retry policy without changing identity

No batch-wide sent=true flag can express that model.

Put template ownership before the queue

The application should own the canonical message data and the selected template version. Render once from an immutable snapshot containing only the fields needed for the alert: order ID, seller display name, item summary, and a stable link target. Store a content hash alongside the version. This costs storage, but it buys a clean answer during an incident: attempt three carried the same seller-visible message as attempt one.

Mustache is one reasonable logic-less syntax for this boundary. Its variables are HTML-escaped by default, sections control repeated or conditional content, and triple mustaches bypass escaping. That last feature deserves a review rule. Seller or buyer input must not reach an unescaped slot.

Provider-owned templates can still be a deliberate choice when non-engineering teams require independent content releases. The trade-off is operational: the delivery record must capture the remote template identifier and revision, and the team must decide whether retry semantics mean "same event" or "same rendered message." If the provider cannot pin a revision, byte-for-byte replay is not a property the application can claim.

Here is the core record and transition logic in Go. The example leaves transport calls behind an interface because queue safety should not depend on one API.

package delivery

import (
    "crypto/sha256"
    "errors"
    "fmt"
)

type State string

const (
    Pending   State = "pending"
    Submitted State = "submitted"
    Delivered State = "delivered"
    Retryable State = "retryable"
    Permanent State = "permanent_failure"
)

type Delivery struct {
    Key             string
    TemplateVersion string
    Payload         []byte
    PayloadHash     [32]byte
    State           State
    Attempt         int
}

func New(orderID, sellerID, channel, version string, payload []byte) Delivery {
    key := fmt.Sprintf("%s:%s:%s", orderID, sellerID, channel)
    copyOfPayload := append([]byte(nil), payload...)
    return Delivery{
        Key:             key,
        TemplateVersion: version,
        Payload:         copyOfPayload,
        PayloadHash:     sha256.Sum256(copyOfPayload),
        State:           Pending,
    }
}

func (d *Delivery) MarkSubmitted() error {
    if d.State != Pending && d.State != Retryable {
        return errors.New("invalid transition to submitted")
    }
    d.Attempt++
    d.State = Submitted
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The database must enforce Key uniqueness. A worker claims a row with a lease, submits the immutable payload, and records the transport's message identifier in the same delivery record. If the process dies in the uncertainty window between remote acceptance and the local commit, the stable key is the only durable way to reconcile instead of blindly duplicating. Where a transport accepts an idempotency key, pass it; where it does not, hold the delivery in an unknown review path until status reconciliation or the retry horizon resolves it.

Recover partial failures without replaying success

First, freeze broad replay controls. Query recipient rows by state and age, not by batch headline. A useful incident view groups counts by channel, template version, state, and last transition time, while retaining the oldest queue age. This exposes a bad template release, a single-channel outage, and worker starvation as different shapes.

Then classify evidence in this order:

  1. Confirm workers are claiming leases and renewing them.
  2. Compare local submitted records with transport message identifiers and later reports.
  3. Move explicit transient outcomes to retryable; move explicit permanent outcomes to permanent_failure.
  4. Requeue only retryable rows whose next-attempt time has arrived.
  5. Keep accepted or delivered recipients out of replay, even when their siblings failed.

Use exponential backoff with jitter and a finite attempt or age limit. The exact numbers belong to the service's delivery objective and transport guidance, not to a copied constant. A new-order alert loses value with time, so the terminal action should usually be an operations signal and a visible undelivered status, not retries that continue after the order workflow has moved on.

Polling also needs restraint. Poll only records that have a transport identifier and lack a terminal result, spread the checks with jitter, and stop at a documented horizon. Callback and poll handlers should call the same compare-and-set transition function. Two ingestion paths are fine. Two state machines are not.

Verify the repair and prepare rollback

Before releasing, run a mixed-outcome test with at least three recipients: one accepted, one transiently rejected, and one permanently rejected. Assert that only the transient row becomes eligible for another attempt and that its payload hash stays unchanged. Deliver the same callback twice and verify one state transition. Deliver an older callback after a terminal result and verify no regression.

Canary the worker change on one queue partition or a small deterministic seller cohort. Watch oldest-ready age, lease-expiration count, attempts by outcome class, time spent in submitted, and terminal failures by channel. The batch completion percentage is useful for a dashboard, but it is too coarse to page or recover from by itself.

Rollback means stopping new claims before changing schema or retry policy, allowing in-flight leases to settle, and restoring the previous worker while preserving delivery rows and immutable payloads. Do not delete the evidence. If a template release is implicated, pin the prior version for new records; existing records should retain their original version unless an operator explicitly creates a corrected notification with a new identity.

Finally, verify sender-side prerequisites independently of queue health. For email, authentication, complaint handling, unsubscribe behavior where applicable, and domain reputation can affect outcomes even when every worker is healthy. Yahoo's sender guidance is a useful primary checklist for those controls. Queue depth cannot diagnose them.

The operating rule

Own enough of the template lifecycle to make replay semantics explicit, and make recipient-channel records the source of truth. Batch for throughput. Diagnose, retry, and audit individually.

This design accepts more rows and more state transitions in exchange for bounded impact: one seller's failure cannot turn into duplicate alerts for everyone else. That is the trade I would take before raising concurrency, shortening poll intervals, or adding another queue.

References

Top comments (0)