Short answer: page on a missed business result, not on a flag changing. For a logistics import, the least complex useful design is an internal control page backed by a small service that can set, list, toggle, and retire flags, while the scheduled worker emits four separate signals: run started, run completed, records accepted, and records rejected. Show the flag revision beside those signals. The on-call can then decide whether to retry an import, revert a flag, or investigate the upstream feed.
The alert arrives at 04:20: manifest_import_results_absent for depot north-17. A useful page contains the last expected window, the last successful result timestamp, accepted and rejected counts, the active flag revision, and the most recent control-plane actor. A page that says only "job failed" is noise with a timestamp attached. Worse, an alert on every flag toggle confuses a control action with customer impact.
The SLO-shaped question is narrow: did the scheduled window produce a valid result before its deadline? Everything else is supporting evidence.
Results first.
What should wake the on-call?
Work backward from the responder's decision. A scheduler heartbeat proves that a process woke up; it does not prove that a manifest was fetched, parsed, accepted, or published. A successful upstream response is similarly incomplete. A zero-record file may be legitimate for one depot and suspicious for another, so a global records == 0 threshold will eventually train the team to ignore pages.
Use four monotonically accumulated event counters, partitioned only by stable operational dimensions such as import type and depot class:
-
import_runs_started_totalidentifies scheduler activity. -
import_runs_completed_totalseparates hung work from work that returned. -
import_records_accepted_totalrepresents usable output. -
import_records_rejected_totalexposes validation failure without pretending rejected rows are success.
Do not put a flag key, user identifier, file name, or manifest ID into metric labels. Those values create an open-ended set of series and make capacity planning depend on business cardinality. Put per-run identifiers in logs or traces, then link them through a generated run ID. The dashboard needs the aggregate for detection and the run ID for diagnosis.
A practical alert evaluates an expected schedule window rather than a fixed process heartbeat. If an import is due every hour and its agreed completion budget is 15 minutes, evaluate whether accepted results advanced inside that window, with a separate warning path for completed runs that rejected every record. Those numbers are examples, not universal defaults; derive the real window from the logistics contract, late-arrival distribution, and error budget.
False certainty is expensive.
How should an internal admin dashboard handle feature flags CRUD?
A simple internal tool still needs a boundary. The browser should not write a shared configuration file or mutate worker memory directly. It calls a control service; that service validates a complete desired state, records a revision, and makes the new value available to workers through a read path. Workers attach the observed revision to each run record. This creates a causal breadcrumb without claiming that every failure after a change was caused by the change.
The following Go types are deliberately boring. Enabled is explicit, retirement is distinct from disabling, and optimistic concurrency prevents two operators from silently overwriting each other.
package flags
import (
"errors"
"time"
)
type Flag struct {
Key string `json:"key"`
Enabled bool `json:"enabled"`
Revision uint64 `json:"revision"`
UpdatedAt time.Time `json:"updated_at"`
UpdatedBy string `json:"updated_by"`
RetiredAt *time.Time `json:"retired_at,omitempty"`
}
type SetRequest struct {
Enabled bool `json:"enabled"`
ExpectedRevision uint64 `json:"expected_revision"`
Reason string `json:"reason"`
}
func (r SetRequest) Validate() error {
if r.ExpectedRevision == 0 {
return errors.New("expected_revision is required")
}
if len(r.Reason) < 8 {
return errors.New("reason must explain the operational intent")
}
return nil
}
Three endpoints are enough: list active and retired flags, put a complete state for one key, and delete one key by retiring it. A toggle button is presentation, not a special storage operation; it reads the current revision and submits the opposite complete state. Treating "toggle" as a blind verb makes retries ambiguous. If the first response is lost, a retry could flip the value back.
Retries happen.
The write handler should authenticate the operator, authorize the flag scope, validate the key against an allowlist, compare the expected revision, persist the new revision and audit record together, and return the stored representation. The worker should continue using its last valid snapshot if a refresh returns malformed data. That is a fail-stable choice, but it needs a staleness signal and a documented maximum age; otherwise a resilient cache becomes an invisible split brain.
Instrument the change, then test the absence
The worker owns result instrumentation because only it knows whether records became usable. The control service owns change events because only it knows who requested the state and which revision won. Join those streams in the investigation view by revision and time, rather than stuffing control metadata into every metric label.
package imports
import (
"context"
"time"
)
type Signals interface {
RunStarted(context.Context, string, uint64)
RunCompleted(context.Context, string, uint64, time.Duration)
RecordsAccepted(context.Context, string, int)
RecordsRejected(context.Context, string, int, string)
}
type RunResult struct {
Accepted int
Rejected int
}
func Execute(ctx context.Context, s Signals, class string, revision uint64, load func(context.Context) (RunResult, error)) error {
started := time.Now()
s.RunStarted(ctx, class, revision)
result, err := load(ctx)
if err != nil {
return err
}
s.RecordsAccepted(ctx, class, result.Accepted)
s.RecordsRejected(ctx, class, result.Rejected, "validation")
s.RunCompleted(ctx, class, revision, time.Since(started))
return nil
}
Notice what this example refuses to do: it does not mark completion on the error path, and it does not convert an error into zero accepted records. Those are different failure modes. The evaluator can distinguish "never finished" from "finished with no usable output," while the on-call can inspect the run log for the exact error. Rejection reason classes must come from a bounded vocabulary; raw error messages belong in logs.
Test absence with a controllable clock. A happy-path test should advance through one schedule window and observe accepted results. Then cover a disabled import, a hung loader, an empty-but-valid source, all rows rejected, a stale flag snapshot, a duplicated write request, and two writers racing on the same revision. The valuable assertion is that the page fires once, at the intended severity, with evidence pointing to a distinct action.
Roll out worker instrumentation before enabling the page. Observe several real schedule cycles, compare the proposed evaluation with expected outcomes, and record which cases would have paged. Deployment order matters: worker emission first, dashboard second, alert last. Reversing it creates an avoidable no-data page.
Buy, build, or keep the first version small?
The decision is less about the CRUD form than the operating surface around it. A platform team should estimate on-call ownership, authentication integration, audit retention, recovery testing, expected flag count, evaluation traffic, and migration cost before debating interface polish.
| Option | Signal-quality fit | On-call load | Lock-in pressure | Capacity question |
|---|---|---|---|---|
| Small internal control service | Exact fit is possible; the team owns correctness | Highest; storage, backup, auth, and alerts remain yours | Low behind a narrow interface | Can cached reads cover peak workers plus admin writes? |
| Shared self-hosted control plane | Common lifecycle behavior can reduce new code | The team patches, scales, restores, and upgrades it | Moderate if its model leaks into workers | What happens to stale reads during control-plane loss? |
| Managed control plane | Local evidence still needs design | Lower infrastructure load; integration ownership remains | Potentially higher through SDK and policy coupling | Is evaluation local or remote, and what is its failure budget? |
For a handful of import gates, start with the smallest service that meets the control requirements and keep evaluation behind an interface owned by the worker. This is not permission to skip backups or authorization. For hundreds of flags, multiple teams, or complex targeting, the homegrown option accumulates policy and lifecycle work quickly; reassess before a basic form becomes critical infrastructure by accident.
This approach has a clear limitation: a small internal service is a poor fit when non-engineers need complex audience targeting, approvals span many organizations, or the platform team cannot own a highly available configuration store. In those conditions, choose a shared self-hosted or managed control plane according to the team's on-call capacity and acceptable coupling, then keep the import-result alert independent of that choice. The trade-off runs in both directions. A broader control plane reduces lifecycle code, but it does not know whether a depot import produced usable manifests; the worker still has to emit that business result. A small service preserves a narrow evaluation contract, but every backup drill, authorization rule, schema migration, stale-cache policy, and audit-retention job lands on the team that built it.
No CRUD screen removes that work.
Capacity planning should include failure mode, not just average request rate. Worker reads may be cacheable and frequent, while admin writes are rare but consequential. Define behavior during store loss, cap snapshot age, measure refresh failures, and prove restoration. A control plane that handles ten times normal traffic but cannot restore revisions is not ready for an on-call dependency.
Deletion is a lifecycle event
A delete button should retire a flag, remove it from normal evaluation, and preserve enough audit history to explain prior behavior. Physical erasure can occur later under a retention policy. This distinction matters because operational evidence and personal data have different reasons for retention. If audit records contain personal data, GDPR Article 17 establishes a right to erasure and lists circumstances in which that right does not apply; legal and security owners must define the applicable policy rather than letting a database default decide it.
Keep actor identity out of metrics. Store it in access-controlled audit records, minimize what is collected, and make retention enforceable. The UI should require confirmation that names the key and current revision, reject stale revisions, and display the retirement result. Recreating a retired key should be deliberate, with a new lifecycle, rather than an accidental effect of a retried request.
Close the loop at the page. A threshold that is too tight wakes someone for routine late arrivals; one that is too loose consumes the logistics error budget before a human can intervene. Review pages by outcome: Was usable data missing? Did the notification arrive early enough to act? Did the evidence identify retry, revert, or upstream investigation? If repeated notifications produce no action, change the signal or move it out of paging. The objective is a small set of trustworthy interruptions tied to scheduled-import results.
Further reading
- GDPR Article 17, "Right to erasure": https://gdpr-info.eu/art-17-gdpr/
Top comments (0)