DEV Community

FletcherVance3712
FletcherVance3712

Posted on

Feature Flag Kill Switch: 4 Signals Before Auto-Disable After Repeated Errors

A healthtech AI agent may retry tool calls, so one failed poll cannot safely mean one failed workflow. Disable too early and an intermittent dependency removes a useful clinical support path; disable too late and the loop can continue accumulating latency, cost, and repeated errors. The choice is therefore a small state machine, not an error counter.

TL;DR: poll an aggregated health snapshot, require four signals—minimum volume, error ratio, consecutive bad windows, and a latency or cost guardrail—and submit an idempotent disable decision with an immutable audit record. Keep evaluation separate from flag mutation. Recovery should be a deliberate, observable transition with a different threshold from shutdown. This makes rollback conservative enough for a clinical workflow while still bounding a runaway agent loop.

How should a feature flag kill switch handle repeated errors?

A raw count discards the denominator. Five errors among six completed agent runs represent a different condition from five errors among fifty thousand, while five identical log lines may all belong to one retrying run. Counting log events also confuses attempts with outcomes, which is precisely the distinction that matters when the agent invokes a slow model or a clinical data tool more than once.

Use a stable unit of work such as agent_run_id, and deduplicate terminal outcomes before aggregation. Each observation should carry a window boundary, cohort or release identity, completed-run count, failed-run count, a latency statistic, estimated cost in the accounting unit chosen by the team, and freshness. The kill decision then has enough context to answer two separate questions: is the feature causing harm, and is the evidence trustworthy?

The four inputs serve different purposes:

  1. Minimum volume prevents a tiny sample from controlling the release.
  2. Error ratio normalizes failures by completed work.
  3. Consecutive bad windows reject a single noisy interval.
  4. A latency or cost guardrail catches loops that remain technically successful while becoming operationally unacceptable.

This mirrors a useful monitoring discipline: observe errors and latency alongside traffic, rather than interpreting one counter in isolation. Saturation may also matter at the dependency boundary, but it should not be smuggled into the same boolean unless the response is genuinely the same. A saturated worker pool may call for admission control; an unsafe agent release calls for rollback.

Keep those responses separate.

Model the decision before the poller

The evaluator should be pure: one snapshot in, one decision out. Poll scheduling, authentication, retries, and mutation belong outside it. That separation matters during an incident because an operator can replay the exact evidence against the exact policy version without contacting the live control plane.

The following Go example uses integers for ratios and cost units, avoiding floating-point comparisons in the control path. Its numbers are illustrative policy inputs, not universal clinical or compliance limits.

package killswitch

import "time"

type Snapshot struct {
    WindowEnd       time.Time
    CompletedRuns   int64
    FailedRuns      int64
    P95LatencyMS    int64
    CostMicrounits  int64
    ConsecutiveBad  int
    ObservedAt      time.Time
}

type Policy struct {
    MinCompletedRuns int64
    MaxErrorBPS      int64 // 1 basis point = 0.01 percentage point.
    BadWindows       int
    MaxP95LatencyMS  int64
    MaxCostMicrounits int64
    MaxSnapshotAge   time.Duration
}

type Decision struct {
    Disable bool
    Reason  string
}

func Evaluate(now time.Time, s Snapshot, p Policy) Decision {
    if s.CompletedRuns < p.MinCompletedRuns {
        return Decision{Reason: "insufficient_volume"}
    }
    if now.Sub(s.ObservedAt) > p.MaxSnapshotAge {
        return Decision{Reason: "stale_snapshot"}
    }

    errorBPS := s.FailedRuns * 10_000 / s.CompletedRuns
    qualityBreach := errorBPS >= p.MaxErrorBPS
    guardrailBreach := s.P95LatencyMS >= p.MaxP95LatencyMS ||
        s.CostMicrounits >= p.MaxCostMicrounits

    if s.ConsecutiveBad >= p.BadWindows && (qualityBreach || guardrailBreach) {
        return Decision{Disable: true, Reason: "repeated_guardrail_breach"}
    }
    return Decision{Reason: "within_policy"}
}
Enter fullscreen mode Exit fullscreen mode

There is an intentional asymmetry here. A stale or undersized snapshot declines to disable rather than treating missing telemetry as proof of failure. Silence is not health, however; the poller must alert separately when snapshots are stale, because a broken measurement path otherwise creates false confidence.

The policy also needs a scope. Prefer the narrowest independently reversible cohort—such as one agent release or one workflow—provided requests can be assigned consistently. A global switch is simpler but expands the blast radius. A cohort switch costs more operational bookkeeping, yet preserves unaffected workflows and produces cleaner rollback evidence. For clinical support systems, that is usually the more defensible trade.

This controller has a firm limitation: it is unsuitable when the feature cannot be isolated from the baseline workflow, when disabling it would remove the only safe clinical path, or when the available telemetry cannot distinguish agent runs from retry attempts. In those cases, choose admission control, a fixed degraded mode, or human approval instead of automatic flag mutation. The trade-off is slower intervention in exchange for avoiding an automated rollback whose effect is broader than its evidence.

Make disabling exactly-once in effect

Networks do not grant exactly-once delivery. The practical target is exactly-once effect: repeated submissions for the same decision converge on one disabled state and one logical audit event. Generate the idempotency key from stable decision identity, for example the feature, cohort, policy version, and breached window end. Do not derive it from the poll attempt time, because every retry would then appear new.

Short timeout. Bounded retry.

The mutation interface can expose compare-and-set semantics without tying the evaluator to a flag vendor. A successful response should distinguish “changed now” from “already disabled,” while both count as convergence. The audit sink should enforce uniqueness on the decision key and retain the evidence used to decide.

package killswitch

import (
    "context"
    "fmt"
)

type DisableCommand struct {
    Feature       string
    Cohort        string
    PolicyVersion string
    WindowUnix    int64
    Reason        string
}

func (c DisableCommand) Key() string {
    return fmt.Sprintf("%s:%s:%s:%d", c.Feature, c.Cohort, c.PolicyVersion, c.WindowUnix)
}

type FlagStore interface {
    DisableIfEnabled(ctx context.Context, feature, cohort, decisionKey string) (changed bool, err error)
}

type AuditSink interface {
    AppendOnce(ctx context.Context, decisionKey string, command DisableCommand) error
}

func Apply(ctx context.Context, flags FlagStore, audit AuditSink, cmd DisableCommand) error {
    key := cmd.Key()
    if _, err := flags.DisableIfEnabled(ctx, cmd.Feature, cmd.Cohort, key); err != nil {
        return fmt.Errorf("disable flag: %w", err)
    }
    if err := audit.AppendOnce(ctx, key, cmd); err != nil {
        return fmt.Errorf("record decision %s: %w", key, err)
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

This compact sample leaves a real transaction boundary visible. If the flag store changes state and the audit write then fails, a retry must repair the missing audit record without toggling the feature back on. In a production design, use a durable decision record and a worker, or an outbox owned by the same database transaction that accepts the decision; then reconcile accepted decisions against both the control state and audit sink. Pretending two independent systems commit atomically would be a more dangerous example.

Audit records should contain evidence, policy version, scope, actor identity, timestamps, prior state, requested state, and mutation result. They should not contain prompts, patient content, or raw tool payloads merely because those fields were available to the poller. Retention and access limits belong in the policy surrounding the sink. The control evidence can be useful without becoming a second clinical-data store.

Alert on the control path, not every retry

Repeated tool-call errors still deserve ordinary telemetry, but paging on each attempt creates noise and makes retries look like separate incidents. Alert when the state machine enters a pending-breach state, when a disable decision is accepted, when mutation cannot converge, and when telemetry becomes stale. Those events correspond to operator decisions.

Keep three identifiers connected across metrics, structured logs, and traces: the agent run, the evaluated window, and the disable decision. The run identifier supports deduplication; the window identifier reconstructs the calculation; the decision key proves that retries converged. Avoid putting high-cardinality patient or prompt values into metric labels. Detailed evidence belongs in access-controlled records linked by opaque identifiers.

Operational status also needs to show the four golden signals at the agent boundary: completed traffic, terminal errors, end-to-end latency, and resource pressure. Cost is an additional guardrail for this workload, not a substitute for those signals. If a loop repeatedly calls a model but eventually returns success, its error ratio can remain green while its latency and cost breach policy.

Rollback must remain available when observability is impaired. Provide an authenticated manual disable path whose action enters the same audit stream, and test it without depending on the poller. Automation narrows response time; it must not become the only route to safety.

Roll out the controller without trusting it blindly

Begin in shadow mode. Evaluate live snapshots, persist proposed decisions, and alert, but do not mutate a flag. Compare each proposal with the underlying run outcomes and confirm that retries collapse into one decision. This phase also reveals whether window boundaries, late-arriving outcomes, and cohort assignment are stable enough for control.

Next, allow automatic disablement for one narrowly scoped cohort while keeping re-enable manual. Exercise four cases in a non-production environment: a single transient error, repeated bad windows above minimum volume, stale telemetry, and a successful disable followed by a lost acknowledgement. The last case verifies idempotency rather than happy-path logic.

Finally, expand scope only after reconciliation shows that every accepted decision has matching control state and audit evidence. Define a recovery policy separately: require healthy windows, an operator approval, or both. Reusing the shutdown threshold for automatic re-enable invites oscillation at the boundary, and a timer alone says nothing about whether the release recovered.

The resulting controller is intentionally conservative. It can stop an unhealthy or runaway clinical agent loop, explain why it acted, and survive duplicate polls, while leaving recovery under an explicit policy. That is the standard a rollback mechanism should meet: bounded harm, reproducible evidence, and no ambiguity about who or what changed production state.

Sources

References:

Top comments (0)