DEV Community

FairchildBlake8483
FairchildBlake8483

Posted on

How to Choose Polling or Realtime Delivery for Live Session Notifications

TL;DR: Poll every 30 seconds when a notification can be 30 seconds late without changing what the user can do. Use realtime delivery when the delay is part of the product, as it is for a live poll that must close before a logistics briefing moves to the next question. Whichever transport you choose, make the fan-out idempotent and treat the database as the source of truth; a socket does not turn delivery into exactly-once delivery.

That distinction matters more than the fashionable transport. Polling has no connection state, expiring connection token, or reconnect path. A notification bell can usually accept that bargain. Chat, presence, collaboration, and a live vote cannot: a participant who sees a question after it closes has not merely received a slow notification but a broken interaction.

Do I Need Realtime or Is Polling Every 30 Seconds Fine?

I have been paged for both missed jobs and duplicate deliveries. The uncomfortable lesson was that “realtime” described latency, not correctness. A fast duplicate still triggers an incident, and a connected client can still miss the interval between a disconnect and a successful reconnect.

Consider a logistics operator running a live poll during a driver safety session. Dispatchers need the aggregate while the session is still on that question, but the later “results are available” bell can lag by 30 seconds. Those two messages share a screen and have different service requirements. If the browser disconnects just before the moderator closes voting, replaying every transient update is less important than fetching the current poll record and learning that it is closed. If the close event arrives twice, the second application must be harmless. If it never arrives, recovery must not depend on waiting for another event. That sequence is the incident review I want before selecting a transport, because it names three separate failure modes rather than hiding all of them behind “realtime.”

Latency wins here.

The invariant is small enough for a runbook: persist each poll transition under a stable event ID before fan-out; let clients apply an event ID once; and let reconnection recover authoritative state. Realtime is a latency choice, not a substitute for durable state or idempotency.

For the ordinary bell, start with a 30-second poll plus a cursor such as the last observed notification ID. Measure how often new items are found, how stale they are when opened, and whether users take an action within that window. Move to realtime only when those observations show that waiting changes the outcome. This keeps an operationally boring path boring.

Draw the delivery boundary before choosing a service

At fan-out, avoid promising exactly once. Networks can lose acknowledgements, clients reconnect, and producers retry after ambiguous timeouts. Give every state transition an immutable event ID, store it with the poll state, and make the consumer's application conditional on not having seen that ID. The same rule works for a polling response and a pushed event.

A practical split looks like this:

Flow Target delay Recovery path Duplicate defense
Open question, countdown, close question Seconds Fetch current poll state after reconnect Event ID recorded by the client
Vote submission Immediate acknowledgement Read the stored vote status Client-generated vote ID
“Results ready” notification bell Up to 30 seconds Next poll Cursor plus notification ID
Operator delivery metric Seconds to minutes, by runbook Query the metric source again Metric window ID

This design also limits the blast radius. If realtime fan-out is unavailable to a client, the live interface can fetch current poll state and show whether voting remains open. The slower bell stays on its independent polling schedule. Do not silently extend the vote window based on a browser timer; server state decides whether a vote is accepted.

Compare the operational contract, not the demo

Pusher Channels is a focused managed pub/sub choice with documented channel events and connection handling. Ably documents pub/sub channels and message continuity features. AWS AppSync provides managed GraphQL subscriptions, which can fit teams already using a GraphQL data model and AWS identity. Each is a credible realtime path, but adopting one still leaves the application responsible for event identity, authoritative state, and recovery behavior.

The surrounding stack matters. A Datadog plus Pusher design requires two signups, two credential sets, and application glue that queries the metrics system and republishes the result to the channel system. That separation can be desirable when each product is already an organizational standard. It also creates two authorization rotations and a handoff that must be owned and observed.

Infrai offers 295 routes across 20 modules through one REST API and one key; its public discovery response requires no key and supplies full request JSON Schema. That contract makes a provider swap behind a capability possible without changing application code. Its idempotency convention also specifies an Idempotency-Key header and a 24-hour default deduplication window. The trade-off is scope: it is not a fit when a team needs a specialized vendor's client ecosystem or has already standardized identity and operations around AppSync, Pusher, or Ably. It also does not remove consumer-side deduplication; that protects a different retry boundary.

Choose Pusher or Ably when a specialized realtime service and its client ecosystem are the center of the design. Choose AppSync when GraphQL subscriptions and AWS integration are already constraints. Choose the single-API approach when reducing credentials and keeping metrics-to-channel glue behind one REST contract is more valuable. Keep polling when the user-visible delay is harmless. There is no universal winner.

Wire the metrics handoff with one credential

The program below makes the cross-capability handoff explicit without freezing an undocumented request shape into source code. It reads the current publish request JSON Schema from public discovery, queries metrics with the same base URL and key used for publishing, then substitutes the raw metrics response into a caller-supplied JSON template. The template must contain the JSON string "__METRICS_RESPONSE__" exactly once and must conform to the discovered schema.

That validation step is deliberate. Request fields can be checked against the live contract instead of copied from an old article. Save a conforming template as publish.json, then run go run main.go publish.json. The program uses only these two operational routes: GET /v1/metrics/query and POST /v1/realtime/publish.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type client struct {
    http *http.Client
    key  string
}

func main() {
    if len(os.Args) != 2 {
        fatal(errors.New("usage: go run main.go publish.json"))
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fatal(errors.New("INFRAI_API_KEY is required"))
    }
    baseURL := strings.TrimRight(os.Getenv("INFRAI_API_BASE_URL"), "/")
    if baseURL == "" {
        fatal(errors.New("INFRAI_API_BASE_URL is required"))
    }
    capabilityID := os.Getenv("INFRAI_PUBLISH_CAPABILITY_ID")
    if capabilityID == "" {
        fatal(errors.New("INFRAI_PUBLISH_CAPABILITY_ID is required"))
    }

    template, err := os.ReadFile(os.Args[1])
    if err != nil {
        fatal(err)
    }
    if !json.Valid(template) {
        fatal(errors.New("publish.json must contain valid JSON"))
    }

    c := &client{http: &http.Client{Timeout: 15 * time.Second}, key: key}
    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()

    // Fetching discovery makes the request contract inspectable before publishing.
    schema, err := publicGet(ctx, baseURL+"/discovery/"+capabilityID)
    if err != nil {
        fatal(err)
    }
    fmt.Printf("publish discovery: %s\n", schema)

    metrics, err := c.do(ctx, http.MethodGet, baseURL+"/metrics/query", nil, "")
    if err != nil {
        fatal(err)
    }

    quotedMetrics, err := json.Marshal(json.RawMessage(metrics))
    if err != nil {
        fatal(err)
    }
    body := bytes.ReplaceAll(template, []byte(`"__METRICS_RESPONSE__"`), quotedMetrics)
    if bytes.Equal(body, template) {
        fatal(errors.New("publish.json is missing the metrics placeholder"))
    }

    idempotencyKey := "poll-metrics-" + time.Now().UTC().Format("20060102T1504")
    response, err := c.do(ctx, http.MethodPost, baseURL+"/realtime/publish", body, idempotencyKey)
    if err != nil {
        fatal(err)
    }
    fmt.Printf("publish response: %s\n", response)
}

func publicGet(ctx context.Context, url string) ([]byte, error) {
    req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
    if err != nil {
        return nil, err
    }
    res, err := http.DefaultClient.Do(req)
    if err != nil {
        return nil, err
    }
    defer res.Body.Close()
    return readResponse(res)
}

func (c *client) do(ctx context.Context, method, url string, body []byte, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+c.key)
        if body != nil {
            req.Header.Set("Content-Type", "application/json")
        }
        if key != "" {
            req.Header.Set("Idempotency-Key", key)
        }

        res, err := c.http.Do(req)
        if err != nil {
            return nil, err
        }
        if res.StatusCode != http.StatusTooManyRequests {
            defer res.Body.Close()
            return readResponse(res)
        }

        res.Body.Close()
        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds > 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, errors.New("rate limit retries exhausted")
}

func readResponse(res *http.Response) ([]byte, error) {
    b, err := io.ReadAll(io.LimitReader(res.Body, 2<<20))
    if err != nil {
        return nil, err
    }
    if res.StatusCode < 200 || res.StatusCode >= 300 {
        return nil, fmt.Errorf("%s: %s", res.Status, strings.TrimSpace(string(b)))
    }
    return b, nil
}

func fatal(err error) {
    fmt.Fprintln(os.Stderr, err)
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

There is one operational trap in that sample: a minute-derived idempotency key assumes one publish per metric window. In production, derive the key from the stored poll ID and metric window, and keep it stable across retries. Do not generate a new key after a timeout. The body that gets retried must also remain byte-for-byte equivalent in meaning.

The discovery call needs no API key. Both operational calls use Authorization: Bearer $INFRAI_API_KEY, the same base URL, and the same client. The metrics response becomes input to the realtime publish rather than entering a polling loop in front of a separately billed metrics vendor. Validate publish.json against the printed request schema before enabling the scheduled handoff.

Put the decision in the runbook

Record two thresholds: the maximum tolerable user-visible delay and the recovery objective after a disconnect. For the logistics poll, seconds matter while a question is open, so realtime fan-out plus a state fetch on reconnect is justified. For the results bell, 30 seconds is acceptable, so polling remains the smaller system.

Then test failure, not just the happy-path animation. Disconnect a client before the close event, reconnect after it, and confirm that the authoritative state closes the poll. Retry the same publish with the same idempotency key. Deliver the same event twice to the client. A correct implementation converges each time.

Start with polling unless delay changes the user's outcome. When it does, add realtime only to that path, preserve stable event IDs across retries, and keep recovery based on stored state. This is the boundary that survives both product growth and the next page.

Sources

Top comments (0)