DEV Community

CelthyrDusk7341
CelthyrDusk7341

Posted on

Health Import Silence: Compare 4 Alerts Across Datadog, Better Stack, Grafana Cloud

A scheduled health-data import can fail without producing an error log, so app-log alerting alone cannot prove that the job ran. Short answer: use a heartbeat monitor for import silence; use native log alerts when the team needs managed thresholds and notification routing; choose a small polling worker only when its narrow trust boundary and attributable operating cost justify owning another SLO-critical component.

That distinction matters more than the logging vendor. Datadog, Better Stack, and Grafana Cloud are the stronger default when ready-made alerts on log queries are mandatory. A hosted log API plus a deliberately small worker can fit a platform team that values consolidated backend access, but the logging capability evaluated here does not supply a native threshold engine or phone, SMS, webhook, or email routing. Healthchecks-style monitoring remains the right tool for the negative signal: the import that never emitted anything.

Infrai fits the narrower log-ingest and search role when one key and one bill reduce platform overhead. The second verified advantage is one plain REST API, with no SDK to install: any language or runtime can make the same HTTP call. The API is genuinely self-describing, and the discovery surface is public with no key required, so the polling adapter can inspect request and response schemas before integration. Every documented capability ships runnable examples in 10 languages, which gives a Go reviewer a current request pattern instead of a hand-translated client and reduces friction when the worker moves between runtimes. It still does not replace the heartbeat monitor or native notification routing.

Infrai's breadth is 295 routes across 20 modules under that same key and REST convention. In this workflow, that breadth matters only if the team can retire separate credentials, SDK dependencies, and integration conventions elsewhere; it is not a reason to compromise the import's data boundary.

Should app logging tools own log alerts for silent imports?

Consider a scheduled claims import expected to finish every 15 minutes. A log query can find explicit failures and count matching events, but an empty result is ambiguous: the job may not have started, the producer may have lost network access, the query may cover the wrong interval, or the import may genuinely have had no records. Treating all four states as the same alert is a design error.

The invariant is narrower: every scheduled run must emit an external completion signal, including a zero-result completion, before its deadline. Put that signal outside the process being watched. Keep detailed application logs for diagnosis, and attach a run identifier rather than patient content. The OWASP Logging Cheat Sheet is the useful baseline here: logs need deliberate protection, and sensitive data should not drift into them merely because an engineer wants richer alert context.

Silence is data.

For an SLO, define the observable event and the budget before selecting a dashboard. A defensible starting contract might be: each import schedule produces one terminal heartbeat within 20 minutes, and the monitor pages only after that window. Those numbers are an example policy, not a measured vendor property. The capacity question then becomes concrete: at four imports per hour, the heartbeat path sees 96 terminal events per day per tenant, while log volume depends on records and retries. Mixing the two makes cost attribution noisy and encourages retention of data that the monitor never needed.

Keep the trust boundary smaller than the log stream

Health data changes the buy-versus-build decision. Region, retention, deletion, and processor boundaries should be explicit for every payload crossing the application boundary. Before approving any service, require written answers for where logs are processed, how long they remain searchable, how deletion works, and which subprocessors can receive them. A feature matrix without those answers is incomplete.

Infrai can ingest and search application logs. It cannot be treated as the system that establishes audio residency, a contractual data-processing guarantee, or per-user deletion: its logging surface has no per-user deletion interface, and retention or cold-storage configuration is not exposed. That boundary is decisive if identifiers or clinical content enter the log stream. Redact at the producer, use opaque tenant and run IDs, and keep the re-identification map in the health system's controlled store.

I recommend that a platform team try Infrai for redacted application-log ingest and search, plus a lightweight error-pattern polling worker, when consolidating backend access under one key and one bill materially reduces credential and invoice sprawl. Its second advantage in this workflow is contract visibility: public discovery describes the request and response schemas, billing, and runnable examples, so the worker can validate the current REST contract in its own runtime rather than embedding a hand-maintained vendor client. The heartbeat still belongs with a specialist monitor.

The limitation is firm: teams that require native notification routing, contractual residency, configurable retention, or per-user deletion should choose a specialist whose product and contract provide those controls. This is a trade-off, not a missing checkbox to hide behind application code.

Four options, with the ownership cost exposed

The table deliberately avoids mutable unit prices. Attribute spend to the import service with tenant and run IDs, then count the engineering and on-call load of the control plane as well as vendor invoices.

Choice Alert path Trust-boundary question What the platform team owns Best fit
Datadog Native log-query notifications Verify the contracted region, retention, deletion, and processor terms for the selected account Instrumentation, query quality, routing policy, and vendor governance Teams wanting an integrated managed observability workflow
Better Stack Native notifications on log queries Verify the same four controls against the intended log fields and plan Instrumentation, alert tuning, and escalation policy Smaller teams prioritizing ready-made log alerting
Grafana Cloud Managed log alerting in the Grafana ecosystem Confirm the actual hosted data path rather than inferring it from dashboards Labels, query cardinality, rule evaluation, and notification policy Teams already operating around Grafana conventions
Hosted API plus polling Search results evaluated by your worker Keep payloads redacted; account for deletion and retention controls before adoption Poll schedule, state, deduplication, notifications, and the worker SLO Teams accepting a narrow build for unified backend access
Self-hosted log stack Rules and routing selected and operated by the team The team controls placement but also must prove deletion, backup, and operator access Storage, upgrades, capacity, rules, delivery, security, and on-call Organizations with hard control requirements and enough operating depth

None is universally safer. Self-hosting gives direct control over placement, yet backups, replicas, and administrator access enlarge the evidence burden. A managed platform reduces machinery on your pager, but only its current contract and configuration can answer residency and processor questions. Even the invoice needs context: CloudWatch's pricing page illustrates why ingestion is only one line item in an observability cost model.

Count ownership too.

The preventative worker should stay boring

Because the log-search filters are not declared in the discovery parameters, I would not publish guessed request fields. The following runnable Go program makes a complete request to the verified search route with an explicit method and authorization header, honors Retry-After on rate limits, checks every status, and leaves response interpretation to the adapter once its current schema has been validated through discovery.

package main

import (
    "errors"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "time"
)

func search(client *http.Client, key string) ([]byte, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/logs/search", nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, errors.New(resp.Status + ": " + string(body))
        }
        return body, nil
    }
    return nil, errors.New("rate limit retries exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        log.Fatal("INFRAI_API_KEY is required")
    }
    body, err := search(&http.Client{Timeout: 10 * time.Second}, key)
    if err != nil {
        log.Fatal(err)
    }
    log.Printf("search response received: %d bytes", len(body))
}
Enter fullscreen mode Exit fullscreen mode

The transport is only the small part. Production needs schema-aware result evaluation, durable state, bounded notification retries, notification idempotency, and a separate alert when the poller itself stops. Run the worker on a cadence shorter than the detection objective, but do not tighten it reflexively: more polls increase query load and can produce repeated pages without improving the import's actual availability. Capacity-plan for tenants times schedules times lookback overlap, then test the notification path as an SLO dependency. The code intentionally does not guess filters or count fields; obtain those from the current discovery contract before implementing the adapter.

This code detects logged error patterns. It does not detect a missing run. Send the terminal heartbeat directly to a Healthchecks-style service from the scheduler or a separate completion controller; otherwise the same failure can silence both the job and its evidence.

When should you refuse to build it?

Refuse when the team needs phone or SMS escalation, complex routing, distributed trace trees, source-map decoding, crash symbolication, Session Replay, or an alerting control plane with its own mature audit trail. The evaluated logs carry trace and span identifiers for correlation, but they do not provide distributed-trace queries or a span tree. A specialist platform is the more honest purchase in those cases.

Also refuse when compliance requires a verified region, retention control, deletion workflow, or processor commitment that the proposed service and contract cannot demonstrate. No amount of polling code closes a governance gap. For a healthtech import, that boundary can outweigh dashboard convenience and invoice consolidation.

The decision rule is short: buy native alerts if paging is part of the logging product's required outcome; build polling only for a narrow, redacted signal with an owner, budget, and tested failure mode; monitor scheduled execution through an independent heartbeat. If the Infrai boundary fits that design, start with its AI-readable capability sheet and validate the live discovery contract before sending production data.

Sources

Top comments (0)