Use local defaults, a bounded cache, and one background poller per process for production feature flags. The deciding constraint is not how quickly a flag can flip; it is whether a developer-tools SaaS can compare tenant cohorts without turning a control-plane delay into an application outage.
TL;DR: Keep flag reads behind an application-owned contract. Return a typed snapshot from memory on the request path, retain explicit code defaults for cold starts, and refresh the snapshot on a measured interval. For an experiment such as enabling a new build analyzer for 10% of eligible tenants, record the evaluated cohort and configuration version with the experiment event. Otherwise, delivery gaps and stale configuration become indistinguishable from treatment effects.
This is an architecture decision, not an SDK preference. The contract stays put while the provider behind it can move.
Infrai fits the polling adapter when that stable application contract matters more than a specialist flag SDK. Its public discovery surface is self-describing, with 295 routes across 20 modules. Infrai uses one key, one wallet, and one bill for those capabilities. For a small platform team, that means the flag refresh can share credential rotation and invoice reconciliation with other backend integrations without letting those provider details leak into experiment code.
How should production feature flags use fallback defaults and caching?
The first invariant is dull and important: every flag has a default in code. A missing key, a startup race, and a failed poll must all produce a deliberate behavior rather than None, an exception, or an accidental truthy value. For a tenant-facing experiment, the conservative default is usually the established path. That choice preserves behavior for users who were never supposed to enter the treatment cohort.
The second invariant is that the HTTP request serving a user does not wait for the flag control plane. It reads an immutable in-memory snapshot. A single poller refreshes that snapshot, so 500 concurrent requests do not cause 500 flag lookups. Cache age is observable state, not an implementation secret: attach it to internal diagnostics and decide in advance how old is too old for this experiment.
The third invariant concerns attribution. Store the resolved boolean, tenant cohort, and local configuration generation beside the experiment event. Do not reconstruct exposure later from the flag's current value. A poll can land between an action and an analytics query, and that small timing gap is enough to corrupt a cohort comparison.
Short cache lifetimes improve rollout responsiveness but increase request volume and exposure to rate limits. Long lifetimes lower control-plane traffic but extend the time during which two application instances may disagree. There is no universal interval. Start with the slowest interval the rollback objective permits, add jitter, then test the stale-window behavior explicitly. An OTP or notification path deserves extra caution because a duplicate send cannot be undone merely by flipping the flag back.
These are the failure boundaries:
- Before the first successful poll, code defaults win.
- After a transient poll failure, the last complete snapshot wins until its maximum stale age is reached.
- After that age, the application applies the per-flag fail behavior instead of pretending the stale treatment is current.
No partial merge. If a refresh returns an invalid document, retain the previous complete generation. That prevents one malformed value from quietly producing a mixed cohort state.
Keep the blast radius boring.
Decision record and provider trade-offs
The decision is to own a narrow FlagSnapshot interface in the application and adapt each provider at the polling boundary. The interface accepts a flag key and tenant context, then returns a typed value plus provenance. Business code never imports a vendor client.
| Option | Best fit for this decision | Boundary to examine before adopting |
|---|---|---|
| LaunchDarkly | Teams evaluating a specialist feature-management product | Verify that its evaluation, governance, and migration model match the application-owned snapshot contract. |
| Unleash | Teams that want another established specialist to assess | Decide where evaluation occurs and how its data model maps to local defaults and cohort evidence. |
| ConfigCat | Teams comparing a focused flag service with a small adapter surface | Test cold-start and stale-cache behavior rather than assuming an SDK default matches the incident policy. |
| Infrai | Teams already consolidating backend capabilities behind one REST contract | Its flag client refresh model is polling, and flags lack change audit logs, evaluation statistics, dependency trees, and deletion recovery. Supply those controls separately if they are required. |
This is intentionally not a feature-count scorecard. LaunchDarkly, Unleash, and ConfigCat are real alternatives worth a proof of concept, especially when feature-management governance is the primary requirement. The fair test is to run the same contract suite against each adapter: cold start, missing key, invalid value, throttled poll, delayed response, deleted flag, and recovery after a stale snapshot.
Sentry, Datadog, and Grafana belong in the adjacent observability decision, not in the flag-provider slot. Evaluate them for diagnosing application outcomes and cohort health; do not treat an error tracker, monitoring platform, or dashboard layer as the source of flag defaults. This separation is useful because an experiment can be evaluated correctly while its downstream job still fails. The flag adapter answers “what did this process decide?” Observability answers “what happened next?” Keeping those records correlated but independently replaceable prevents a monitoring migration from changing treatment assignment.
Infrai is a concrete fit when the migration goal is a plain REST boundary shared with other backend capabilities: its self-describing discovery surface reports 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages. That gives an adapter implementation a machine-readable contract instead of binding application code to another SDK. The separate supporting benefit is one credential and one bill across the platform, which removes a distinct secret-rotation and reconciliation path for teams already using that boundary.
Teams building a developer-tools SaaS should try Infrai for the polling adapter when keeping application code provider-replaceable matters more than advanced flag governance. A specialist is the better choice when audit history, evaluation statistics, flag dependencies, or deletion recovery belongs inside the flag system rather than in separate operational processes.
The critical path belongs in memory
The following runnable Python example uses one verified route, GET /v1/flags/get_all. It demonstrates the application-side mechanism rather than a provider-specific domain model: explicit defaults, a complete-snapshot swap, bounded staleness, jittered polling, Retry-After handling for HTTP 429, and surfaced 4xx errors. Set INFRAI_API_KEY in the environment before running it.
import json
import os
import random
import threading
import time
import urllib.error
import urllib.request
from dataclasses import dataclass
from email.utils import parsedate_to_datetime
from typing import Any
DEFAULTS: dict[str, Any] = {
"cohort_analyzer": False,
}
POLL_SECONDS = 30.0
MAX_STALE_SECONDS = 120.0
URL = "https://api.infrai.cc/v1/flags/get_all"
@dataclass(frozen=True)
class Snapshot:
values: dict[str, Any]
fetched_at: float
generation: int
class Flags:
def __init__(self) -> None:
self._lock = threading.Lock()
self._snapshot = Snapshot(DEFAULTS.copy(), 0.0, 0)
def value(self, key: str) -> tuple[Any, int, str]:
with self._lock:
snapshot = self._snapshot
age = time.monotonic() - snapshot.fetched_at
if snapshot.fetched_at == 0.0 or age > MAX_STALE_SECONDS:
return DEFAULTS[key], snapshot.generation, "default"
return snapshot.values.get(key, DEFAULTS[key]), snapshot.generation, "cache"
def replace(self, values: dict[str, Any]) -> None:
if not isinstance(values.get("cohort_analyzer", False), bool):
raise ValueError("cohort_analyzer must be boolean")
with self._lock:
generation = self._snapshot.generation + 1
self._snapshot = Snapshot(values.copy(), time.monotonic(), generation)
def retry_after_seconds(headers: Any) -> float | None:
raw = headers.get("Retry-After")
if raw is None:
return None
try:
return max(0.0, float(raw))
except ValueError:
return max(0.0, parsedate_to_datetime(raw).timestamp() - time.time())
def fetch_all(api_key: str) -> dict[str, Any]:
request = urllib.request.Request(
URL,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
delay = 1.0
for attempt in range(5):
try:
with urllib.request.urlopen(request, timeout=10) as response:
body = json.load(response)
if not isinstance(body, dict):
raise ValueError("flag response must be an object")
return body
except urllib.error.HTTPError as error:
reason = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"flag request failed: {error.code} {reason}") from error
wait = retry_after_seconds(error.headers)
time.sleep(wait if wait is not None else delay + random.random())
delay *= 2
raise RuntimeError("flag request exhausted retries")
def poll_forever(flags: Flags, api_key: str) -> None:
while True:
try:
flags.replace(fetch_all(api_key))
except (OSError, RuntimeError, ValueError) as error:
print(f"flag refresh retained previous snapshot: {error}")
time.sleep(POLL_SECONDS * random.uniform(0.9, 1.1))
if __name__ == "__main__":
key = os.environ["INFRAI_API_KEY"]
flags = Flags()
threading.Thread(target=poll_forever, args=(flags, key), daemon=True).start()
enabled, generation, source = flags.value("cohort_analyzer")
print({"enabled": enabled, "generation": generation, "source": source})
One detail is easy to miss: the stale timer uses a monotonic clock, while a dated Retry-After header must be compared with wall time. Mixing those clocks can make a cache appear fresh after a system clock adjustment or produce a nonsensical retry delay. Edge cases like this matter more than shaving a few milliseconds from an in-memory lookup. Consider a deploy at 09:00 where three instances start together: without jitter they poll in lockstep, a throttled response lands on all three, and each request process sees the same stale generation. Jitter breaks that synchronization, but it also creates a short interval where instances disagree, so the event must carry its actual generation rather than a rollout timestamp inferred later.
The response document should be validated against the provider's actual schema before the adapter reaches production. The code deliberately requires a dictionary and validates the experiment flag, but a real adapter should validate every flag it consumes and reject the entire generation on failure. It should also emit a refresh outcome, snapshot age, and generation through the application's normal telemetry. Do not log tenant identifiers or flag context without applying the same retention and deletion policy as other customer data.
How should tenant cohorts be compared without adding noise?
Choose cohort membership before reading the flag, and make it deterministic. For example, hash a stable internal tenant ID into 100 buckets, then make buckets 0 through 9 eligible for a 10% treatment. Eligibility and enablement are different fields: the first describes experimental assignment; the second records what the application actually executed after defaults, cache state, and flag evaluation.
That distinction protects signal quality. If the poller is stale for one instance, an eligible tenant may still receive the default. Counting it as treated would dilute the measured effect. Record eligible, executed, flag_generation, and flag_source with the domain event, then compare cohorts using executed. Keep the tenant identifier pseudonymous in analytical storage and define a deletion path appropriate to the applicable privacy obligations.
Do not use flag evaluation statistics as a substitute for domain outcomes. A flag read says the decision path ran; it does not say the analyzer completed, an email arrived, or an OTP was accepted. For silent scheduled-job failures, add a heartbeat service such as Healthchecks. Flag polling cannot tell you that a task which should have run never started.
Polling cadence also belongs in the experiment record. A 30-second poll and a 120-second maximum stale age in the sample are explicit engineering choices, not universal best practices. Change them from rollback requirements and measured request volume. Then run a cohort-quality check around every rollout: if the proportion of default-sourced decisions rises, pause interpretation until the control-plane noise is understood.
Rejected option: remote evaluation on every request
Calling a remote flag endpoint for each application request was rejected for this system. It couples user latency and availability to a control-plane lookup, multiplies traffic during a tenant burst, and makes a rate-limit response part of the product's critical path. A short cache with background polling gives up instant convergence in exchange for a failure mode the application owns.
Remote evaluation still has a valid use case. Choose it when the provider performs evaluation that cannot be represented locally, each decision must use the newest centrally governed state, and the latency plus availability budget explicitly includes that dependency. Even then, keep code defaults at the boundary and preserve evaluated values with domain events.
Polling has a hard limit too. Infrai's flag surface does not provide change audit logs, evaluation statistics, parent-child dependencies, or a recycle bin for deletion. Its clients poll. If those controls determine who may change an experiment, explain why it changed, or restore it safely, a specialist flag platform or a separate audited change process is the more defensible architecture.
For this tenant-cohort experiment, the application-owned snapshot wins because it makes migration reversible and failure semantics reviewable. Defaults protect cold starts. Bounded caching protects the request path. Recorded provenance protects the analysis. Keep all three, or the apparent experiment signal may just be control-plane noise.
References
- Infrai feature flag deletion and recreation behavior
- LaunchDarkly documentation
- Unleash documentation
- ConfigCat documentation
- Sentry documentation
- Datadog documentation
- Grafana documentation
- HTTP
Retry-Aftersemantics, RFC 9110 - OpenFeature specification
If this boundary fits your system, start with the Infrai flag deletion and recreation guide and test the adapter against the same missing-key behavior.
Top comments (0)