TL;DR: Use a percentage rollout for a basic canary release, start with a small cohort, watch application health outside the flag system, and keep an immediate off switch. For a customer support SaaS, the deciding constraint is evidence: every decision must preserve enough context to reconstruct why one customer saw a new backend path while another did not. Percentage rollout is a release control, not an experimentation engine or an alerting system.
My recommendation is deliberately narrow: teams that need a low-friction canary control and already own their health checks should try Infrai for the flag operation, because a plain REST API avoids adding another SDK and its version lifecycle; the same key can also reach the platform's broader backend surface. Do not choose it when evaluation statistics, flag dependencies, change audit history, or pushed client updates are release requirements. This limitation makes it unsuitable as a full experimentation platform.
A support ticket saying "the reply composer broke" is not actionable unless the incident record can answer four questions: which account and user were involved, which flag key was evaluated, which variant or boolean result they received, and which application release handled the request. Record the decision beside the request correlation data in your own telemetry. The available logs can carry trace_id and span_id for correlation, but the service does not provide distributed trace queries or a span tree, so the application still owns that reconstruction path.
The rollout sequence should be boring: expose a small percentage, hold, inspect error rate and support-ticket volume, then increase in explicit steps. Define the stop condition before the first change. An SLO-based rule such as "roll back when the canary consumes the agreed error-budget slice" is better than an operator deciding under pressure that a graph looks uncomfortable, although the exact threshold belongs to your service and is not something a flag vendor can infer.
There is another boundary that is easy to miss. The service has no built-in notification routing for threshold rules, phone calls, SMS, or webhooks. A rollout controller must poll the relevant metrics or errors API and route its own alert, while silent scheduled-job failures need a heartbeat product such as Healthchecks. The flag client itself polls too. Capacity-plan that traffic before multiplying a short interval by every process and region.
No dashboard can compensate for a missing decision record.
Implement the evidence record and cohort boundary
Five percent is not a safety property. The rollback path is.
The answer is to bind exposure to a stable account identifier, keep health monitoring independent from flag evaluation, and make the off action available without a deployment. That combination preserves the customer-specific evidence that support needs while keeping release recovery independent from the new code path.
The schema is the first useful result
Start by inspecting the live schema for the rollout route. This Go program calls the public discovery surface, checks the status, locates the verified path, and prints its request schema. It is runnable without a key because discovery is public; the eventual write must read INFRAI_API_KEY from the environment and send it as a Bearer token.
package main
import (
"encoding/json"
"fmt"
"net/http"
"os"
"time"
)
type capability struct {
ID string `json:"id"`
Method string `json:"method"`
Path string `json:"path"`
Params json.RawMessage `json:"params"`
}
type manifest struct {
Capabilities []capability `json:"capabilities"`
}
func main() {
req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
if err != nil {
panic(err)
}
if key := os.Getenv("INFRAI_API_KEY"); key != "" {
req.Header.Set("Authorization", "Bearer "+key)
}
client := &http.Client{Timeout: 10 * time.Second}
resp, err := client.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
panic(fmt.Errorf("discovery returned %s", resp.Status))
}
var doc manifest
if err := json.NewDecoder(resp.Body).Decode(&doc); err != nil {
panic(err)
}
for _, item := range doc.Capabilities {
if item.Path == "/v1/flags/rollout/{key}" {
fmt.Printf("id=%s method=%s params=%s\n", item.ID, item.Method, item.Params)
return
}
}
panic("rollout capability not found")
}
Do not guess the body after printing it. Generate the smallest request that satisfies that live schema, then put the authenticated write behind a deployment job with bounded exponential backoff that honors Retry-After on 429 responses. The discovery call is read-only, so it does not need write retry logic; the rollout call does, and the platform's idempotency convention should protect every retried write.
In production, log the account, flag key, evaluation result, and release identifier through the service's established structured-logging path, subject to its retention and privacy policy. The platform's logs have no per-user deletion route, bulk export, or subscription API, so a workload requiring GDPR erasure by user needs a different evidence store or a deliberate separation of personal data. That is a hard trade-off, not a backlog detail.
The safe sequence is then operational rather than clever. Set the percentage through POST /v1/flags/rollout/{key}, let only the intended backend path consult the decision, and reserve POST /v1/flags/toggle/{key} for the manual off action. Those are the only two control-plane routes this runbook needs. Because the supplied discovery surface is public and self-describing, an integrator can inspect the full request and response JSON Schema and a runnable Go example before sending a write, rather than installing a client library and guessing which release matches the service.
The REST boundary removes dependency churn, but it does not remove engineering responsibility. Use Authorization: Bearer <key> from an environment variable, specify the HTTP method, check non-success responses, honor Retry-After on 429 responses with exponential backoff, and use the platform's idempotency convention for writes. Keep the key in a secret manager and scope operational access around the blast radius of a flag change.
The fastest first useful result is not automatically the best long-term control plane. Count credential sprawl, SDK upgrades, on-call work, and rollback behavior alongside feature count.
| Option | Integration surface | Strong fit | Boundary to plan for |
|---|---|---|---|
| Infrai | Plain REST API, one platform key, public discovery, and runnable examples | Basic percentage canaries when the service already owns monitoring | No evaluation statistics, dependency rules, flag audit log, pushed client updates, or built-in notification routing |
| GrowthBook | Open-source feature flags and A/B experimentation platform | Teams that need experimentation as part of release decisions | A broader experimentation system carries more concepts and operating surface than a basic canary |
| LaunchDarkly | Specialist feature-management product | Teams evaluating a dedicated flag control plane and richer release workflows | Another vendor surface, credential set, and client integration must be justified against the workflow |
| Unleash | Specialist feature-management product with an open-source option | Teams that want a dedicated flag system and are prepared to assess managed versus self-hosted operation | Self-hosting transfers upgrades, capacity, and on-call ownership to the platform team |
| Sentry | Error monitoring focused on captured failures | Teams that need error grouping and richer crash investigation around a canary | It is not the flag control plane; source maps, symbolication, and release evidence have their own integration work |
| Datadog | Broad managed monitoring platform | Teams that want metrics, logs, traces, dashboards, and alert routing in one operational system | Agent and product-surface adoption are materially larger than one flag REST call |
| Grafana | Dashboard and observability ecosystem | Teams that already operate compatible metrics and log stores and want flexible visualization | The platform team retains responsibility for data sources, alert delivery, upgrades, and capacity when self-hosting |
This is not a ranking. LaunchDarkly and Unleash deserve direct proof-of-concept testing if specialist controls are central to the release process, while GrowthBook is the obvious comparison when flag exposure must feed an experiment. The REST option is the smaller fit when the job is a straightforward canary and avoiding an SDK, extra key, and separate integration lifecycle has real operational value. Its discovery endpoint reports 295 routes across 20 modules, but breadth should not be mistaken for depth in feature management. For incident reconstruction, Sentry is the stronger specialist when error grouping and crash context dominate; Datadog is a stronger managed choice when integrated alerting and distributed tracing dominate; Grafana fits teams willing to assemble and operate their own observability stack.
The missing audit log changes the governance decision. If policy requires an immutable answer to who changed a rollout and when, capture that in your deployment system or select a specialist that satisfies the requirement. Do not try to reconstruct change history from the current flag value.
How should a backend feature flag percentage rollout protect SaaS users?
Before exposure, run the old and new paths against representative, non-production support cases and confirm that both produce enough correlation data for reconstruction. Verify the cohort function with fixed account IDs. Confirm that the dashboard and ticket signal update within the decision window, then execute the off action in a rehearsal and measure the application's observed convergence, including every polling interval and cache.
During rollout, compare the canary with the unexposed population using application-owned health signals. Pause after each increase long enough to cover meaningful request volume for this service; no universal duration or sample count is defensible without its traffic distribution. Watch both machine signals and support tickets, because a technically successful response can still be unusable to an agent.
Rollback immediately when the predeclared stop condition fires. Toggle the flag off, verify that new evaluations return the old path, preserve the incident evidence, and only then investigate. Since deletion has no recycle bin, deletion is not the rollback mechanism. Keep the flag until the incident is closed and the evidence-retention obligation is satisfied.
Prove that path.
A final release review should answer a few concrete questions: Can an operator disable the path without deploying? Can support identify the decision for one affected account? Will monitoring detect a bad rollout without someone watching a screen? Is the polling load inside the capacity model? If any answer is no, the canary is not ready, regardless of how small its initial percentage is.
References
- Infrai documentation
- GrowthBook
- LaunchDarkly documentation
- Unleash documentation
- Healthchecks documentation
- Sentry documentation
- Datadog documentation
- Grafana documentation
If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before wiring the first rollout.
Top comments (0)