TL;DR: For a logistics notification service, page from a bounded failure-rate counter, preserve structured logs for diagnosis, and attach every signal to a release cohort. Do not make repeated log searches the primary detector. A cheap alert that cannot distinguish the current release from the previous one is expensive during rollback, because it can tell you that deliveries are failing without telling you whether reverting code will stop them.
The governing trade-off is detection detail versus rollback certainty. Logs carry the detail; counters produce a stable alerting signal; a release identifier connects both to the action an operator can take. Build that small path first. Buy broader retention, search, correlation, or operating capacity only when the on-call load and scale justify transferring those responsibilities.
Should cheap log-based failure alerts search error status or poll metrics?
Imagine a release exercise for a service that accepts notification jobs and hands them to a downstream delivery provider. Version release-b begins returning more final failures than release-a, while both versions serve traffic during a gradual rollout. A search for status=error answers one narrow question: error records exist. It does not, by itself, prove that release-b caused the increase, that the failures are final rather than retryable, or that rolling back will repair deliveries already rejected for bad recipient data.
That distinction is the incident lesson. An actionable release alert needs a denominator, a failure class, and a release cohort. Without all three, the page is a prompt to investigate rather than evidence for rollback. The numerator should count terminal outcomes that the service owns or can classify, while the denominator counts eligible delivery attempts under the same dimensions; the cohort identifies deployed behavior, not a host that may outlive several releases. Keep the dimensions bounded: channel, outcome class, and release are useful; recipient, tracking number, and message ID belong in logs, not in the alerting counter. This is the trade-off I would make: accept less detail in the paging signal so its capacity and meaning remain predictable, then retain job-level detail in the diagnostic event where an operator can use it after the page.
Short version: cardinality is capacity.
Rollback comes first.
A notification can fail in several operationally different ways. Input can be invalid before any provider call. A provider can reject a request permanently. A transient dependency error can enter a retry path. A worker can stop consuming. A request can be accepted while the final delivery result arrives later. Combining those states under one error token creates a large number with very little rollback value, so define the state machine before defining the alert.
For this service, I would use a deliberately small outcome vocabulary such as accepted, retryable_failure, terminal_failure, and delivered, then document which transition increments which counter. That is a design proposal, not a claim that every provider exposes the same lifecycle. If the system cannot observe final delivery, name the signal submission_failure rather than implying that it measures delivery.
The preventative path is small enough to review
The following Go example shows the boundary. It emits a structured event for diagnosis and updates an abstract counter for alerting. The event includes a pseudonymous job identifier, while the counter excludes it. Both carry the same release and outcome classification, which makes an alert comparable with the diagnostic records without turning identifiers into metric dimensions.
package notify
import (
"context"
"encoding/json"
"io"
)
type Counter interface {
Add(ctx context.Context, name string, value int64, attributes map[string]string)
}
type Attempt struct {
JobID string
Channel string
Release string
}
type Result struct {
Outcome string
Reason string
}
type failureEvent struct {
Event string `json:"event"`
JobID string `json:"job_id"`
Channel string `json:"channel"`
Release string `json:"release"`
Outcome string `json:"outcome"`
Reason string `json:"reason"`
}
func RecordResult(ctx context.Context, out io.Writer, counters Counter, a Attempt, r Result) error {
attrs := map[string]string{
"channel": a.Channel,
"release": a.Release,
"outcome": r.Outcome,
}
counters.Add(ctx, "notification_attempts", 1, attrs)
if r.Outcome != "retryable_failure" && r.Outcome != "terminal_failure" {
return nil
}
event := failureEvent{
Event: "notification_attempt_finished",
JobID: a.JobID,
Channel: a.Channel,
Release: a.Release,
Outcome: r.Outcome,
Reason: r.Reason,
}
return json.NewEncoder(out).Encode(event)
}
Production code also needs an explicit policy for failures in the telemetry path. Losing a log write should not silently change a delivery result, yet blocking the request indefinitely to preserve a diagnostic event is usually the wrong availability trade. Buffer limits, backpressure, drop accounting, and shutdown flushing therefore belong in the design review. The exact policy depends on whether the service handles synchronous requests, asynchronous jobs, or both; the important part is that telemetry failure behavior is chosen, tested, and visible.
The alert then evaluates a ratio over a window, separated by release and failure class. Its policy can be stated without binding the system to a query language: page only when terminal failures divided by eligible attempts exceed the service's chosen threshold, the denominator exceeds a minimum traffic floor, and the condition persists for the chosen number of evaluation windows. Compare the candidate release with the stable cohort before rollback. The threshold and windows must come from the service objective and traffic shape, not from a copied dashboard.
No magic number fits here. A five-minute window may be empty for one channel and hide a fast regression in another. Capacity planning starts with observed attempts per channel and release, then asks how many evaluation periods are needed to avoid reacting to one failure in a nearly idle cohort. Test that arithmetic with sparse traffic, a traffic spike, and a complete worker stall.
Search, poll, or evaluate counters?
These are different mechanisms, even if one backend can perform all three.
| Mechanism | Best role | Rollback weakness | Operating burden |
|---|---|---|---|
| Scheduled log search | Retrospective diagnosis and low-volume audits | Parsing changes, late records, and missing denominators can change the apparent rate | Retention, indexing, query scheduling, and schema control |
| Counter evaluation | Primary rate and availability alerts | Bounded dimensions omit per-job detail | Instrumentation, aggregation, and evaluation state |
| Status polling | Checking a genuinely asynchronous final state | Poller lag or failure becomes part of detection latency | Scheduling, cursor state, rate limits, and duplicate handling |
Use polling only when the source of truth changes after the initial request and no dependable event arrives at that boundary. Polling a metric merely to discover whether its threshold was crossed recreates an alert evaluator in application code, including missing-data semantics, retries, deduplication, and notification delivery. That may be acceptable for a tiny system, but it is still software with an on-call owner.
Log search remains valuable. During an incident, an operator can pivot from release-b plus terminal_failure to the reason and pseudonymous job ID, inspect representative failures, and decide whether the release changed behavior. The same event can support a periodic audit that compares counted outcomes with logged outcomes. It should not be the sole fast path unless the team has deliberately accepted the query scheduler and indexing path as part of alert availability.
There is another blind spot: no delivery attempts means no failure records. A queue can stop draining while a log query reports zero errors. Pair the failure-rate alert with a progress signal such as queue age or time since the last completed attempt, selected according to the actual job model. This is why an alerting design cannot begin and end with searching for HTTP error status.
Silence can be failure.
Rollback safety changes the buy-versus-build boundary
A platform decision should account for the service hidden behind the invoice: ingestion durability, query and evaluation behavior, retention, access control, upgrades, and the hours consumed during an incident. Price alone is a weak axis. On-call load and the ability to leave are usually harder constraints.
| Decision area | Build or self-host when | Use a managed service when | Exit evidence to retain |
|---|---|---|---|
| Collection | The team already operates a shared telemetry path | Maintaining ingestion would compete with delivery reliability work | Versioned event schema and portable emission code |
| Alert evaluation | Requirements are few, bounded, and covered by existing operations | Evaluation availability and notification routing need a dedicated owner | Alert definitions, test fixtures, and SLO rationale |
| Search and retention | Volume is controlled and operators can own storage lifecycle | Retention, indexing, and access operations dominate team time | Exportable records and documented field semantics |
| Incident workflow | Existing paging and runbooks cover the path | Routing, deduplication, and audit needs exceed local capacity | Provider-independent severity and ownership model |
I would reject any option, managed or self-hosted, that cannot preserve release as a queryable dimension from page to investigation. I would also budget storage and evaluation load from event volume, retention, dimension counts, and peak ingestion rather than average request rate. A cheap pilot with unbounded identifiers can become an operational liability before it becomes a financial one.
Feature flags deserve special care. Martin Fowler's discussion of feature toggles distinguishes categories with different lifetimes and dynamics, and it also describes cohort-based release techniques. For alerting, record the stable configuration or cohort identifier that can explain behavior, but do not attach an arbitrary set of flag names and user-level assignments to every counter. Logs can retain the richer decision context. Counters need a bounded release comparison.
The safest rollback is also tested before production. Run the candidate and stable cohorts together, inject a terminal failure into the candidate path, verify that only its numerator changes, and confirm that the page links to records carrying the same release value. Then roll the candidate back and confirm that new attempts shift to the stable cohort while older asynchronous work remains attributable to the code that processed it. Do not relabel old work after deployment; that destroys the comparison the operator needs.
When this design does not apply
Rate alerts are the wrong primary detector when a single failure is catastrophic, because waiting for a denominator and several windows delays action. They are also insufficient when the application sees only submission acceptance and another system owns final delivery. In that case, establish an observable handoff and assign the end-to-end objective to the boundary that can actually see it.
Very low traffic needs different treatment. A ratio can swing from zero to 100 percent on one attempt, so a synthetic transaction, a maximum queue-age check, or a direct terminal-event alert may carry more information than a percentage. Keep the structured event either way; it is the evidence that explains the page.
The final rule is narrow: page on a bounded service-level signal, diagnose with structured events, and make release identity survive the entire path. If an alert cannot support a rollback decision, label it investigative and keep it out of the urgent paging route. That standard controls noise, clarifies ownership, and leaves room to change the storage or evaluation implementation later.
Top comments (0)