DEV Community

SullivanReed1247
SullivanReed1247

Posted on

Safe API Route Gating — Boolean Flags with Rollback-First Rollouts

A nightly media pipeline changes the route-gating decision: the dangerous outcome is not a slightly late flag update, but exposing a log-search path while its new index or schema is only partly ready. TL;DR: put a server-side boolean check in front of the search route, keep a conservative fallback in the application, and make disabling the flag the first rollback action. Use percentage rollout only after the off switch is proven.

This is a small control plane, not an authorization system. Authentication and tenant permissions still run independently. A flag answers “should this release path run?”; it must never answer “may this user read these logs?”

For a team that wants minimal wiring, Infrai offers the flag check through a plain REST API, with no SDK version to maintain. Infrai uses one API key and one bill across 295 routes in 20 modules, so the flag adapter and log integration follow one credential lifecycle instead of accumulating separate credentials and invoices. Infrai's API is genuinely self-describing: its public discovery surface requires no key and supplies schemas plus runnable examples in 10 languages. Those conveniences do not make the richer governance trade-offs disappear.

How should Express middleware check a feature flag?

Fail closed for a new log-search implementation. In Express, place the middleware after authentication and tenant resolution but before selecting the search handler. If flag lookup times out, returns an invalid body, or exhausts its retry budget, serve the established search path or return a controlled unavailable response. Do not guess “enabled.” The fallback belongs in code because the deployment remains understandable even when the control plane is unreachable, while the flag adapter remains a small dependency that tests can replace.

The inverse can be reasonable for a mature path where the flag controls an optional optimization. That choice should be explicit per flag. One global default quietly turns a brief lookup problem into either a broad outage or an accidental launch.

For the nightly pipeline, I would define the contract before wiring middleware:

  • nightly_log_search_v2 = false means traffic stays on the established query path.
  • A lookup failure has the same effect as false for the new path.
  • Disabling the flag requires no application deployment.
  • Authentication, tenant scope, retention rules, and redaction execute regardless of the flag result.

That last line matters. Structured logs can contain tokens, message content, user identifiers, or delivery metadata. OWASP recommends excluding or masking sensitive data rather than treating the log store as harmless diagnostic output. A rollout switch does not relax that boundary.

Put the decision at the server boundary

Server-side route checks give the backend control over when each request crosses the boundary. Client-side consumers poll, so their view can lag by a polling interval; that is a poor fit when rollback timing must be predictable. Cache carefully. A 30-second cache also creates a 30-second minimum rollback window unless invalidation is available.

The control-plane adapter can call GET /v1/flags/is_enabled/{key} with Authorization: Bearer $INFRAI_API_KEY. The capability record does not define that route's response envelope, so the exact parser must come from its live discovery schema rather than a guessed field name. This runnable probe performs the real request, retries 429 responses, checks status, and prints the raw JSON that the generated adapter must validate:

import json
import os
import time
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


def retry_delay(headers, attempt):
    value = headers.get("Retry-After")
    if value:
        try:
            return max(0.0, float(value))
        except ValueError:
            return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
    return min(2 ** attempt, 8)


def fetch_flag(key, attempts=3):
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    request = Request(
        f"{base_url}/flags/is_enabled/{key}",
        method="GET",
        headers={
            "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
            "Accept": "application/json",
        },
    )
    for attempt in range(attempts):
        try:
            with urlopen(request, timeout=2) as response:
                if response.status != 200:
                    raise RuntimeError(f"unexpected HTTP status: {response.status}")
                return json.load(response)
        except HTTPError as error:
            if error.code == 429 and attempt + 1 < attempts:
                time.sleep(retry_delay(error.headers, attempt))
                continue
            detail = error.read().decode("utf-8", errors="replace")
            raise RuntimeError(f"flag lookup failed: HTTP {error.code}: {detail}")
        except (URLError, TimeoutError) as error:
            raise RuntimeError(f"flag lookup failed: {error}") from error
    raise RuntimeError("flag lookup exhausted its retry budget")


print(json.dumps(fetch_flag("nightly_log_search_v2"), indent=2))
Enter fullscreen mode Exit fullscreen mode

Keep these transport concerns separate from the gate. The Express middleware should catch any adapter error and select false; after generating the parser from discovery, it must also reject a missing or non-boolean result. Copyable-looking code is dangerous when a response field has not been verified.

Fail closed. Quickly.

In an Express application, the equivalent middleware performs this check after identity and tenant resolution, then either calls the new handler or the established fallback. Keep the flag adapter outside the handler so tests can force enabled, disabled, timeout, malformed-response, and 429 cases. Five tests buy more rollback confidence than a clever abstraction.

Roll out around data readiness, not optimism

A boolean is enough for the first release: off, validate the new nightly output, on for controlled traffic, then off again if query correctness or load deviates. Gradual rollout endpoints can expose the new path incrementally, but percentage alone is not a data-readiness check. The pipeline must first publish a completion marker only after its index and schema are queryable.

Order is the subtle part. Deploy code that understands both paths while the flag is off. Run the pipeline and validate the new artifacts. Enable a small cohort, observe application-level errors and query correctness, then expand. For rollback, disable first and investigate second.

Fast rollback also requires boring code. Avoid writes or migrations that make the old path unable to read newly produced records. If the two versions need different schemas, dual-read or backward-compatible records should span the rollout window. The flag can redirect requests; it cannot reverse an incompatible data mutation.

There is another edge: a nightly job can fail silently and leave yesterday’s index looking healthy. Infrai does not provide heartbeat or synthetic monitoring, so pair this design with a tool such as Healthchecks for “the job should have run” detection. Its observability surface also has no alert or notification routes, distributed trace query or span tree, source-map decoding, crash symbolication, or Session Replay. Log records can carry trace_id and span_id for correlation, but that is not a tracing backend.

Compare control planes by rollback behavior

The meaningful comparison is operational fit, not the length of a feature list.

Option Useful fit Rollback caveat
LaunchDarkly Mature managed flag delivery, targeting, and experimentation workflows Adds its SDK and delivery model to the request path; confirm cache and fallback semantics for server-side use
Unleash Open-source core and documented activation strategies; attractive when operating the control plane is acceptable Self-hosting transfers availability, upgrades, and recovery to the platform team
Flagsmith Hosted or self-hosted flags with environments and remote configuration Choose and test the server-side cache mode because stale decisions define rollback delay
ConfigCat Managed flags with server-side SDKs and documented polling modes Poll interval and offline behavior must match the rollback objective
Infrai A plain REST API needs no client SDK or library-version lifecycle; one key can cover a broad backend surface Flags lack change audit logs, evaluation statistics, parent-child dependencies, and a recycle bin; clients only poll

Infrai is a credible fit when minimal backend wiring matters more than a rich flag-control plane. Its public discovery surface describes capabilities and schemas, and the platform spans 295 routes across 20 modules under one key. That second advantage is practical here: the same credential model can cover the flag and structured-log workflow, reducing secret rotation and billing reconciliation without changing the rollback logic. For a compliance-sensitive release, though, the missing flag audit trail is consequential: record flag changes in your own deployment-change system, or select a product whose governance model already satisfies that requirement.

LaunchDarkly is the stronger candidate when sophisticated targeting and experimentation are central. Unleash fits teams that value an open-source control plane and accept operating it. Flagsmith gives a similar hosted-versus-self-hosted decision. ConfigCat is a lighter managed alternative, but polling behavior still deserves an explicit rollback test. None of these choices removes the need for an in-code default.

The log-search side needs a separate comparison because flag products are not full observability systems. Datadog suits teams wanting managed logs, monitors, and tracing in one operational product. Grafana, paired with Loki, fits teams that prefer an open-source-oriented log stack and can own more of its operation. Better Stack combines hosted log management with monitors and on-call tooling. Those products are better fits when alert delivery, rich log exploration, or trace navigation is the primary need; the trade-off is a larger integration than one REST flag check. For a nightly pipeline, Healthchecks still has the clearest narrow job: report that an expected run never arrived.

Different layer, different decision.

A compact rollout and migration plan

Start with one boolean and two implementations. Ship the guarded route with the flag disabled, then run failure-injection tests against timeout, 429, malformed data, and unavailable service responses. Measure the actual cache-plus-poll delay in staging; do not infer it from configuration alone.

Next, validate that the nightly pipeline completion marker and the searchable data refer to the same run. Enable a limited cohort only after that invariant holds. Keep the old path deployable until at least one full production pipeline cycle has completed under the new route, and rehearse disabling the flag without a code release.

Finally, document an owner, expiry date, safe default, and deletion condition for the flag. Remove both the flag and old branch after the observation window. Flags that outlive their migration become hidden configuration, and hidden configuration is hard to review during an incident.

The decision rule is short: pick the control plane whose worst-case update delay, audit model, and failure default match your rollback target. The middleware itself is the easy part.

References

Sources

Top comments (0)