DEV Community

marcorossi4891
marcorossi4891

Posted on

Frontend Feature Flags: Backend API Polling for Silent Import Detection

Short answer: expose one server-evaluated control document, poll it with bounded jitter, and use its cohort policy to decide when a scheduled import is late. Keep rollout assignment, environment rules, and alert state on the backend. The browser may display status, but it must not become the authority for either rollout eligibility or paging.

For an edtech import pipeline, the useful invariant is precise: every enabled school cohort should produce a successful import result within its declared schedule plus a grace window. A flag changes which cohorts are observed or how many consecutive late evaluations are required; it does not erase history. That separation gives operators a reversible way to tune signal quality without turning every delayed file into an alert.

This is an architecture decision record for that boundary. The hard part is not putting a Boolean in a component. It is keeping a gradual operational change consistent when tabs sleep, requests overlap, deployments move independently, and a missing result could mean an outage or an expected school-calendar gap.

What must remain true when the control changes?

The first invariant is ownership. The backend evaluates environment, cohort, rollout percentage, and emergency disablement from trusted attributes. A browser receives a small, non-secret decision document. It never receives targeting rules that expose school identifiers, internal segments, or compliance-sensitive attributes.

The second invariant is monotonic observation. Suppose the current control requires three consecutive late evaluations before an alert opens. Lowering that value to two may allow the next evaluation to open an alert, but toggling the feature off must not delete the accumulated observations. Otherwise, an operator can hide a continuing delivery gap by changing configuration. Retain observations with their timestamps and policy version; suppress notification separately.

The third invariant is a bounded freshness contract. Each response carries a version and an expiry time. The client may keep the last valid document until expiry. After expiry it enters a declared stale state rather than pretending that the cached decision is current. A stale flag is an observable state, not a Boolean value.

The failure boundary follows. Configuration storage and assignment are backend concerns. Fetch scheduling and rendering are frontend concerns. Import-result detection belongs to the monitoring worker, while notification deduplication belongs to durable alert state.

Decision options and their noise costs

The choice is less about request syntax than about behavior during partial failure.

Option Source of truth Failure behavior Signal-quality consequence Appropriate use
Build-time environment toggle Deployment artifact Fixed until another build Stable, but slow for incident-time adjustment Policy that should change only with a release
Browser-evaluated rules Each active tab Tabs can disagree due to stale state or clock differences Duplicate or inconsistent status Cosmetic experiments using public attributes
Backend-evaluated document with polling Server plus bounded client cache Last valid value works until explicit expiry Consistent assignment and measurable staleness, with polling traffic Reversible operational controls
Streaming updates Long-lived connection Reconnect and resume become correctness concerns Lower propagation delay, but more connection-state noise Controls requiring latency below the poll interval

For scheduled imports, the backend-evaluated document is the balanced choice. A control change taking tens of seconds to arrive may be acceptable when the underlying cadence is measured in minutes or hours. That is a design assumption, not a universal fact: record a propagation objective and choose the poll interval from it.

Streaming remains valid when an emergency control must propagate faster. It is rejected here because connection lifecycle, proxy timeouts, resume tokens, and fan-out health add states to observe without improving detection of a slowly scheduled process.

How should a React frontend poll a backend for feature flags?

A production design needs a control read and an ingestion path for import outcomes. Avoid making the browser infer pipeline health from loosely timed endpoints. The control response should be cacheable but revalidated, carry a schema version, and distinguish document identity from its validity window.

The following Python expresses the backend decision. The hash provides deterministic assignment; it is not a security primitive. Use a stable opaque tenant key, never an email address or student identifier.

from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
import hashlib


@dataclass(frozen=True)
class ControlDocument:
    schema_version: int
    policy_version: str
    enabled: bool
    late_after_seconds: int
    consecutive_misses: int
    expires_at: str


def bucket(opaque_school_key: str, rollout_salt: str) -> int:
    material = f"{rollout_salt}:{opaque_school_key}".encode("utf-8")
    digest = hashlib.sha256(material).digest()
    return int.from_bytes(digest[:8], "big") % 10_000


def evaluate_control(*, opaque_school_key: str, environment: str,
                     rollout_basis_points: int, rollout_salt: str,
                     now: datetime) -> ControlDocument:
    assigned = bucket(opaque_school_key, rollout_salt) < rollout_basis_points
    return ControlDocument(
        schema_version=1,
        policy_version="import-silence-v3",
        enabled=environment == "production" and assigned,
        late_after_seconds=7_200,
        consecutive_misses=3,
        expires_at=(now.astimezone(timezone.utc) + timedelta(minutes=2)).isoformat(),
    )
Enter fullscreen mode Exit fullscreen mode

A value of 2,500 basis points represents 25% of 10,000 deterministic buckets. Keep the salt stable while increasing exposure so an existing cohort remains assigned. Changing both percentage and salt reshuffles the population and ruins comparison between stages.

The frontend loop should revalidate, reject malformed or expired data, and notify subscribers only when the document changes. In a React or Next.js application, put this behind a shared store so ten mounted components do not create ten timers. This Python specifies the behavior without coupling the decision to a frontend library.

import asyncio
from datetime import datetime, timezone
import random


async def poll_controls(fetch, publish, interval_seconds=30.0):
    etag = None
    last_valid = None
    while True:
        status, next_etag, payload = await fetch(etag)
        if status == 200 and payload and payload.get("schema_version") == 1:
            etag, last_valid = next_etag, payload
            publish({"state": "current", "control": payload})
        elif status == 304 and last_valid:
            publish({"state": "current", "control": last_valid})
        elif last_valid is None:
            publish({"state": "unavailable"})

        if last_valid:
            expiry = datetime.fromisoformat(last_valid["expires_at"])
            if datetime.now(timezone.utc) >= expiry:
                publish({"state": "stale", "control": last_valid})

        await asyncio.sleep(interval_seconds * random.uniform(0.8, 1.2))
Enter fullscreen mode Exit fullscreen mode

Conditional requests reduce transfer when nothing changes. Jitter spreads requests instead of aligning every tab on a 30-second boundary. Pause ordinary polling while a page is hidden, then revalidate when visible. The monitoring worker does not pause; browser visibility must never govern alert correctness.

One race deserves explicit handling. A slow earlier request can arrive after a newer policy. Compare a server-issued version before publishing and discard regressions. Do not use the browser clock to order versions.

This approach has a real limitation: polling exchanges propagation speed for simpler failure handling. A 30-second interval can leave a tab behind for nearly that long, plus network delay, and shortening the interval raises request volume across every open tab. It is also not appropriate for a safety control that must change almost immediately. Streaming or a server-side enforcement point wins in that case. For an import monitor, the trade-off is acceptable only when the documented propagation objective is longer than the worst expected poll delay and the backend continues enforcing the actual alert decision.

How does the monitor distinguish silence from an expected gap?

A timestamp cannot answer that alone. The monitor needs a schedule record, a last successful result, and the policy used for each evaluation. Holidays, district calendars, maintenance holds, and onboarding affect whether silence is exceptional. Model them as eligibility states rather than adding ever-larger grace periods.

For each eligible cohort, calculate due_at from the schedule and then late_at = due_at + grace_window. Increment the miss counter only when an evaluation occurs after late_at and no qualifying result has arrived. Reset it on a qualifying result, not on an API read or flag refresh. Open one alert when the threshold is crossed, and attach later evaluations to it until recovery.

Quiet matters.

One alert. One owner.

Emit structured events for policy evaluation, late detection, state transition, and notification outcome. RFC 5424 defines severity values numerically from Emergency at 0 through Debug at 7, with lower values representing greater severity. Map an opened alert to the locally agreed severity; do not label every late evaluation as an error. A miss below the threshold is often informational telemetry. The transition deserves the louder signal.

Useful fields include an opaque cohort key, scheduled time, last-success time, policy version, freshness state, miss count, and transition reason. Avoid student data and raw import contents. This compliance boundary also helps debugging: responders can explain why an alert opened without seeing the payload.

Test the state machine with result-before-deadline, exact-boundary, three-misses-then-recovery, disabled-with-history, expired-document, and mid-evaluation-policy-change cases. During deployment, first record shadow decisions without notifications. Compare would-open transitions with known outcomes before enabling a small stable cohort.

Rejected option and the case where it wins

We rejected a build-time environment variable as the primary control. It couples alert tuning to deployment, cannot express a stable percentage cohort without extra machinery, and gives operators no independent way to stop notifications while preserving detection state.

It still has a valid use. A build-time switch fits a compile-time capability, a legally constrained regional bundle, or a feature whose activation must pass through the same review and artifact promotion as code. Its slowness is then governance, not a defect.

Advance the rollout only when control freshness meets its objective, shadow decisions match the scheduled-import model, and the ratio of actionable transitions to suppressed late evaluations remains acceptable to the on-call team. Roll back notification eligibility without deleting observations. This preserves evidence, limits noise, and keeps the UI honest about stale state.

References

Top comments (0)