A feature flag kill switch must remain usable during an outage, but health monitoring cannot stop at process uptime when a healthtech pricing change may be charging the wrong amount. Treat every flag decision as an observable business event, then make containment depend on both service health and pricing invariants.
Short answer: put the new rule behind a server-side kill switch, record every evaluation with a stable rollout and rule revision, and monitor outcome ratios alongside latency and errors. During an outage, freeze further exposure first, disable the rule through a pretested control path, and preserve the decision trail. A fast switch without reconstructable evidence stops new damage but leaves the incident team guessing about which invoices, notices, or accounts require repair.
What must a feature flag kill switch prove during an outage?
It must prove more than "the flag is off." For each pricing decision, an operator should be able to answer which rule revision ran, which cohort was evaluated, what amount category resulted, and when the decision occurred. Healthtech adds a sharp boundary: observability records should use opaque account and request identifiers rather than names, email addresses, phone numbers, or clinical details.
The useful unit is a decision envelope. It is emitted at the same boundary that chooses the rule, not later in an HTTP access log where the business context has already disappeared. Logs should be treated as event streams, which also means the application should write the event and leave routing and storage to the execution environment.
No payload dump.
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
import hashlib
import json
import logging
logger = logging.getLogger("pricing.decisions")
@dataclass(frozen=True)
class PricingDecision:
occurred_at: str
request_id: str
account_key: str
rollout_id: str
rule_revision: str
flag_enabled: bool
result: str
def opaque_key(account_id: str, audit_salt: str) -> str:
value = f"{audit_salt}:{account_id}".encode("utf-8")
return hashlib.sha256(value).hexdigest()[:20]
def emit_decision(event: PricingDecision) -> None:
logger.info(json.dumps(asdict(event), separators=(",", ":")))
event = PricingDecision(
occurred_at=datetime.now(timezone.utc).isoformat(),
request_id="req_7f91d2",
account_key=opaque_key("acct_internal_42", "rotate-with-policy"),
rollout_id="pricing_rollout_06",
rule_revision="rule_17",
flag_enabled=True,
result="new_rule_applied",
)
emit_decision(event)
The result field is deliberately categorical. Do not copy plan descriptions, medical context, raw request bodies, or calculated totals into a general-purpose log merely because they would make debugging convenient. Keep the minimum fields needed to join a pricing decision to the restricted system of record under the organization's retention and access rules.
Step 1: Separate control health from business health
An endpoint returning success does not show that a pricing rollout is correct. Four signals cover different failure classes: evaluation availability, request failure rate, latency, and the distribution of pricing outcomes. The last signal catches a valid response that applies a new rule far more often than the rollout policy intended.
Use counts for the numerator and denominator. A lone percentage is hard to interpret when traffic is low, delayed, or uneven across cohorts.
| Signal | Question it answers | Containment clue |
|---|---|---|
| Evaluation count by revision | Which logic actually ran? | A disabled revision still receiving evaluations |
| Eligible and applied counts | Did exposure match policy? | Applied outcomes exceed the eligible cohort |
| Error and latency series | Is the request path degrading? | A change aligned with rollout exposure |
| Missing-decision count | Are requests bypassing instrumentation? | Outcomes exist without a decision envelope |
Do not collapse these into one green light. A control plane can be reachable while cached flag state is stale; the application can be fast while its cohort predicate is wrong. Conversely, a downstream dependency may raise latency even though the new rule is innocent. Separate signals prevent an automatic response from turning correlation into a verdict. The explicit tradeoff is speed against confidence: broad automation reduces reaction time, but a business-invariant trigger needs a narrow, well-understood action. I would allow a breached exposure invariant to freeze the next cohort automatically because that action adds no new accounts; I would reserve full disablement for a high-confidence condition defined before launch or an operator who can inspect the evidence. That choice contains uncertainty instead of pretending the first correlated graph has already established cause.
Step 2: Make the emergency path boring
The switch must sit before new-rule calculation and side effects. If the code creates an invoice, sends a price-change email, and only then checks the flag, disabling it is theater. For email and SMS flows, include the rule revision in the internal message job metadata so queued communications can be reconciled without putting sensitive recipient data into the decision stream.
from dataclasses import dataclass
from decimal import Decimal
from typing import Callable
@dataclass(frozen=True)
class Quote:
amount: Decimal
rule_revision: str
def quote_price(
base_amount: Decimal,
new_rule_enabled: bool,
apply_new_rule: Callable[[Decimal], Decimal],
) -> Quote:
if not new_rule_enabled:
return Quote(amount=base_amount, rule_revision="baseline_12")
return Quote(
amount=apply_new_rule(base_amount),
rule_revision="rule_17",
)
Keep the fallback deterministic and already deployed. The outage is the wrong time to discover that the baseline branch no longer accepts the current data shape. Test both branches on every release, including the boundary where a request begins under one flag state and a retry arrives after the switch changes.
Retries are the nasty edge. A stable idempotency key should identify the commercial operation, while the decision event records each attempt and its rule revision. The repair query can then distinguish a retry from a second purchase without guessing from timestamps.
One key, many attempts.
Step 3: Rehearse the decision, not just the toggle
A useful drill starts with a small synthetic cohort and injects a violated invariant. Confirm that the alert contains the rollout ID, affected revision, first observed time, and links to the relevant internal runbook and dashboards. Then freeze expansion. One operator disables the new rule; another verifies that new decision events show the baseline revision and that queued customer communications are held or reconciled according to policy.
Fast is measurable, but avoid inventing a universal target. Set an internal objective for detection-to-freeze and freeze-to-verification, record both timestamps, and revise the objective after drills. A tiny service with an on-call generalist has a different safe response window from a staffed operations team.
The reconstruction query should return three sets: operations definitely evaluated by the suspect revision, operations definitely evaluated by the baseline, and operations with missing or ambiguous evidence. That third set matters most. Silence is not proof of safety.
Consider a hypothetical SaaS with 10,000 eligible accounts and a first checkpoint set to 5%, or 500 accounts. The monitoring view should reconcile those 500 eligibility decisions against the number of new-rule outcomes before anyone expands to 25%. If it shows 520 applied outcomes, the arithmetic is already enough to freeze expansion even when latency and error graphs look normal. If it shows only 470 decision events but 500 completed pricing operations, the 30-record gap becomes its own incident set; operators should not classify those accounts as baseline or new-rule exposure until the restricted system of record resolves them. This example is intentionally count-based. Ratios alone would hide whether a dramatic-looking change came from five operations or five hundred, and an aggregate total would not identify the accounts that need review.
import sqlite3
def affected_operations(
connection: sqlite3.Connection,
rollout_id: str,
suspect_revision: str,
) -> list[tuple[str, str, str]]:
rows = connection.execute(
"""
SELECT request_id, account_key, occurred_at
FROM pricing_decisions
WHERE rollout_id = ? AND rule_revision = ?
ORDER BY occurred_at
""",
(rollout_id, suspect_revision),
)
return list(rows)
Store the event stream in a system where retention, access, and deletion behavior match the organization's obligations. More data is not automatically better evidence. Log ingestion can carry volume-based costs, so bounded fields and explicit sampling rules matter, but cost should never justify sampling away the very decisions needed to establish impact.
Step 4: Roll out with reconstruction checkpoints
Before any exposure, publish the rollout ID and rule revision, exercise the baseline branch, and verify that synthetic decisions appear in the health view. Start with a cohort small enough to inspect. Expand only after the team can reconcile eligible, evaluated, and applied counts for the completed checkpoint.
If a signal crosses its predeclared boundary, stop expansion. If the evidence implicates the new rule, disable it and verify fresh baseline decisions. Preserve the suspect interval and affected identifiers before changing dashboards or retention settings. Customer correction, compliance review, and communication happen from that frozen evidence set, not from an improvised search assembled hours later.
For an existing boolean flag, migration can stay compact: add a rollout ID and rule revision to the evaluation event, instrument applied and missing-decision counts, deploy and test the dormant baseline path, then run a synthetic containment drill. Only after those checks should the new pricing cohort receive traffic.
The result is a kill switch that supports two jobs under pressure: containing new exposure and explaining past exposure. Both are required for a pricing rollout that may touch invoices and regulated customer communications.
Contain first. Reconstruct next.
References
- The Twelve-Factor App, "Logs": https://12factor.net/logs
- Amazon Web Services, "Amazon CloudWatch Pricing": https://aws.amazon.com/cloudwatch/pricing/
Top comments (0)