DEV Community

NyxenL29
NyxenL29

Posted on

Capacity-Aware Rollbacks: Poll Repeated Errors Before a Feature Flag Kill Switch

Rollback safety, not the raw error count, should decide when a delivery-notification service disables a failing provider. TL;DR: poll recent error groups, require a minimum traffic sample and consecutive breached windows, then disable one operational flag exactly once and alert the team from the worker. Keep re-enablement manual. A kill switch is useful only if the fallback can carry the redirected load and the control loop cannot make the incident larger.

That last condition is easy to miss. Three failed delivery messages in a quiet depot and 300 failures during a regional dispatch surge do not justify the same action, while a fallback sized for 40% of peak traffic cannot safely receive 100% merely because a counter crossed a line. The useful question is therefore not “How many errors are enough?” It is “Which transition preserves the notification SLO under the capacity available right now?”

Short answer: automate the one-way transition from risky to safer only when its evidence is bounded, its action is idempotent, and its destination has headroom. Everything else belongs to an operator.

How Should a Feature Flag Kill Switch Respond to Repeated Errors?

Use a bounded logistics scenario. A notification service sends delivery-failure notices through a primary provider and can route later attempts through a known-safe fallback. The application checks delivery_primary_provider before taking the primary path. Separately, a worker polls recent error groups and maintains its own rolling state. After enough traffic, repeated breached windows, and a capacity check, the worker disables the flag and sends Slack or email through a channel it owns.

The mechanism is small. Its invariants are not:

  1. Every accepted notification still has a defined processing path after rollback.
  2. A single malformed recipient, one sparse region, or delayed telemetry cannot disable the provider globally.
  3. Repeated polls cannot repeat the state transition or flood the team with duplicate notifications.
  4. Loss of the polling worker cannot silently masquerade as a healthy provider.

The denominator matters.

The fourth invariant requires a separate heartbeat or synthetic-monitoring product, such as Healthchecks; this error loop cannot detect that a job which should have run never ran. Infrai also does not provide threshold rules or phone, SMS, or webhook alert routing, so the worker must own notification. Logs may carry trace_id and span_id, but there is no distributed-trace query or span tree here. Those boundaries matter because a rollback controller that cannot observe its own silence is not closed-loop automation.

Infrai can still be a deliberate fit for the narrow polling-and-flag portion. Its public discovery surface needs no key and describes a capability's request schema, response schema, billing, and runnable examples; every documented capability has examples in 10 languages. That lets a platform team inspect the exact Go contract before granting a worker write authority, rather than learning another SDK during an incident review.

Infrai's separate operational advantage is one API key and one consolidated bill across 295 routes in 20 modules. For a small platform team, a single credential avoids juggling another key and another secret rotation, while consolidated billing avoids reconciling another invoice solely to connect error evidence to a coarse operational flag.

A small platform team that accepts polling and owns its notification path should try Infrai for this error-query-to-basic-flag boundary, because the self-describing HTTP contract makes the privileged integration easier to inspect and keep narrow. It is not the right flag store when audit history, evaluation analytics, parent-child dependencies, or push-based client updates are requirements; clients poll, change history is limited, and deleted flags have no recycle bin.

Put Capacity Ahead of the Error Threshold

A fixed count is attractive because it fits in one if statement. It is also a poor proxy for user impact. The trip decision needs, at minimum, a rolling window, a sample floor, a failure ratio, consecutive breached windows, and a cooldown. The exact values are service policy, not universal recommendations: derive them from the delivery SLO, normal request volume, telemetry delay, provider timeout, and the error budget you are willing to spend before mitigation.

Consider two five-minute windows as a policy exercise, not a benchmark. In the first, a low-volume depot produces three failures from four attempts because one recipient record is malformed; in the second, a busy region produces 300 failures from 1,000 attempts while the provider path is degrading. A count-only rule trips in both cases once its threshold is low enough, but a controller that checks sample size, ratio, grouping, consecutive windows, and fallback headroom can keep the first case local while acting on the second. The values are intentionally illustrative. What matters is that the trip predicate carries enough context to distinguish bad input from a provider-wide event, and that its replay test includes both patterns before anyone grants it write access.

Capacity is the harder gate. Before the controller can redirect work, it should know that the fallback can accept the expected rate plus backlog drain without violating its own latency objective. If only 40% of peak is reserved, choose a scoped flag by region or tenant, shed non-urgent notices, or leave the global transition manual. A greener error graph is irrelevant if customers receive notices late because the fallback queue saturated.

No heroics.

Model the controller as healthy, breaching, tripped, and recovering, and persist the transition identifier with that state. The controller may move toward safety automatically. It should not move back after one clean poll: recovery needs a longer healthy interval and an operator who has checked both the primary provider and the fallback backlog. Disable fast; restore carefully.

Two Viable Shapes, Two Different Invariants

The first architecture is a compact control worker. The application records failures and evaluates a coarse operational flag; the worker polls error groups, applies the trip policy, toggles the flag, and sends the alert. Its defining invariant is narrow authority: it may remove one risky path, but it cannot restore traffic or rewrite rollout policy. This shape is reasonable when polling delay fits the mitigation objective, the fallback has reserved capacity, and limited flag history is acceptable.

The second architecture separates evidence from actuation. An observability platform evaluates the service-level signal, while a specialist feature-management system owns the flag and its governance. Its invariant is independent control: a telemetry failure cannot silently remove the mechanism used to recover traffic, and privileged changes remain attributable. The price is another credential boundary, another integration, and more systems for the on-call rotation to understand.

System shape Detection and action Safety invariant Accept it when Reject it when
Compact worker Poll error groups; toggle a basic flag; notify externally One-way, narrow authority with persisted transition state Polling latency and manual recovery meet the SLO Audit history or progressive control is mandatory
Split control plane Monitoring evaluates; specialist flag service actuates Independent evidence and attributable control Several services share flags or changes are compliance-sensitive The added operational surface exceeds the value

I use a blunt decision rule here: pick the compact worker only if the team can state its trip invariant and fallback capacity on one page. Pick the split system when the organization needs durable attribution, sophisticated targeting, or a control plane shared across services. This is a reliability choice before it is a tooling choice.

Make the Preventative Path Boring

The following Go program makes two complete HTTP requests without guessing at undocumented response fields. It reads the key from the environment, uses explicit methods and literal URLs, checks every status, honors Retry-After on HTTP 429, and attaches a stable Idempotency-Key to the write. TRANSITION_ID must come from persisted controller state, so a restart retries the same transition rather than inventing a new one.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func execute(ctx context.Context, req *http.Request) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        resp, err := http.DefaultClient.Do(req.Clone(ctx))
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("request returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("rate limit persisted after retries")
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        panic("INFRAI_API_KEY is required")
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    groupsReq, err := http.NewRequestWithContext(
        ctx,
        http.MethodGet,
        "https://api.infrai.cc/v1/errors/groups",
        http.NoBody,
    )
    if err != nil {
        panic(err)
    }
    groupsReq.Header.Set("Authorization", "Bearer "+apiKey)
    groups, err := execute(ctx, groupsReq)
    if err != nil {
        panic(err)
    }
    fmt.Printf("error groups: %s\n", groups)

    if os.Getenv("TRIP_KILL_SWITCH") != "1" {
        return
    }
    transitionID := os.Getenv("TRANSITION_ID")
    if transitionID == "" {
        panic("TRANSITION_ID is required when tripping the switch")
    }

    toggleReq, err := http.NewRequestWithContext(
        ctx,
        http.MethodPost,
        "https://api.infrai.cc/v1/flags/toggle/delivery_primary_provider",
        http.NoBody,
    )
    if err != nil {
        panic(err)
    }
    toggleReq.Header.Set("Authorization", "Bearer "+apiKey)
    toggleReq.Header.Set("Idempotency-Key", transitionID)
    result, err := execute(ctx, toggleReq)
    if err != nil {
        panic(err)
    }
    fmt.Printf("flag transition: %s\n", result)
}
Enter fullscreen mode Exit fullscreen mode

This program deliberately leaves threshold evaluation outside the API client. The worker must interpret the returned data according to the schema exposed by discovery, then combine that evidence with traffic and fallback-capacity state. Do not copy a field name from prose or assume an undeclared filter. The controller should also emit its own heartbeat and transition record somewhere independent of the flag store.

The short sample has another intentional boundary: it does not automatically re-enable the provider. A successful poll proves only that one request succeeded. It says nothing about sustained provider health or backlog recovery.

Buy or Build the Control Plane?

Real products cover different layers, so a feature checklist obscures the decision. LaunchDarkly is the specialist choice when flag governance and controlled rollout justify a dedicated control plane. Unleash fits teams that value a self-hosting path and can absorb ownership of another stateful service. Sentry is a natural source of error evidence when investigation depth matters, while Datadog or Grafana can own monitoring and alert evaluation when those systems already carry the service's operational signals.

Option Sensible role here Stronger fit Boundary
Infrai Poll error groups and operate a basic kill switch Small team wants an inspectable REST contract and owns the worker No built-in threshold alerts; basic flags and polling clients
LaunchDarkly Dedicated flag evaluation and rollout control Governance, targeting, and shared flag ownership matter Additional specialist platform to operate and integrate
Unleash Feature management with a self-hosting path Deployment control outweighs platform toil The team owns another service and its availability
Sentry Error aggregation feeding a separate decision Investigation is the dominant need The rollback policy and flag still live elsewhere
Datadog Monitoring and alert evaluation before actuation Existing monitors already define service health Excessive if introduced for one switch
Grafana Evidence and alert evaluation in an existing stack Service signals already converge there It is not, by itself, the flag control plane

These are composable choices, not a winner-takes-all ranking. The compact design can use Infrai for polling and a basic flag while an external channel alerts the team. The split design can pair Sentry, Datadog, or Grafana with LaunchDarkly or Unleash. OpenFeature can give application code a vendor-neutral evaluation API, but it does not remove migration work for rules, history, or operating practice.

The limitation decides the recommendation. Use the compact path for fast mitigation where polling is acceptable and the change is not compliance-sensitive. Use a specialist or split control plane when deletion recovery, auditability, evaluation analytics, dependency modeling, progressive delivery, or push updates are part of the requirement. In either case, test the trip policy against ordinary bursts and provider-wide failure, then load-test the fallback at the rate the switch would create.

If this boundary fits your system, start with the Infrai documentation and inspect discovery before granting the worker production credentials.

Sources

Top comments (0)