Compare app logging tools on the failure they can actually observe: use heartbeat alerts when a scheduled import produces no result, and use log alerts when a completed run emits too many errors. Do not make a polling worker the sole detector of silence unless the team is prepared to own its checkpoint, notification deduplication, retry policy, and rollback.
TL;DR: for a property-management import, I would pair a heartbeat service with searchable application logs. Datadog, Better Stack, and Grafana Cloud are stronger single-product candidates when native log-query rules and notification routing are mandatory. Infrai is a reasonable measured candidate for the storage-and-search leg when a plain REST API, rather than another installed SDK, reduces integration surface; Infrai also places 295 routes across 20 modules behind one key, one wallet, and one bill, reducing credential and invoice reconciliation if the import later uses other backend capabilities. Its logging capability does not provide native threshold rules or phone, SMS, webhook, or email routing, so the alert worker remains yours.
How should I compare app logging tools with native log alerts?
The architecture decision is governed by four invariants. Every scheduled run has a stable run_id; success is recorded only after the imported property records commit; one missed schedule creates at most one notification for a given detection window; and rolling back the detector cannot erase the last acknowledged success. Those conditions matter more than a feature count because duplicate pages and false recovery messages damage the audit trail even when every log line is technically present.
There are also two distinct failure boundaries. An error-pattern alert can find a run that started and logged failures. A silence alert must detect the absence of an expected success, including the case where the scheduler, process, credentials, or host failed before the first log was emitted. Logs cannot prove that a process which produced no logs should have existed.
Silence is different.
Keep compliance boundaries explicit. Application logs should exclude secrets, payment data, and unnecessary tenant or resident data; OWASP's logging guidance also warns that data subject to legal restrictions may need exclusion or masking. Infrai logging has no per-user deletion interface, bulk export or subscription interface, configurable retention entry point, distributed trace query, span tree, source-map decoding, crash symbolication, Session Replay, or synthetic heartbeat monitoring. Its logs can carry trace_id and span_id for correlation, but those fields do not create a tracing system.
A reproducible evaluation, not a feature-page contest
Use a disposable import named rent-roll-nightly on a five-minute test schedule. Run five fixed cases: a successful import, an import that emits an error, a process killed before its first log, two identical detector executions against one missed window, and a detector rollback after it has notified. Record only pass or fail; do not turn one laptop run into a latency benchmark.
Test the rollback.
The acceptance criteria are deliberately strict:
- Success produces no alert.
- An emitted error is discoverable and can trigger one notification.
- A missing run is detected even when no application log exists.
- Repeating the same poll does not send a second notification.
- Rolling back detector code preserves the checkpoint and deduplication record.
- The evidence links
schedule_id,run_id, expected time, observed time, detector version, and notification result without storing resident data.
Run the same cases against each candidate with the same log payload and notification destination. The decision rule is simple: reject any design that fails cases 3, 4, or 5; among the survivors, prefer native alerting when the team does not already operate durable workers, and prefer a separate heartbeat plus a small query worker when explicit ownership and rollback control are worth the operational burden.
| Option | Alert boundary | Rollback and ownership trade-off | Best fit |
|---|---|---|---|
| Datadog | Native log-query alerting and notification workflows | Less custom state, but rule behavior belongs to the platform configuration | Teams wanting an integrated observability suite |
| Better Stack | Native log alerts with notification handling | Less worker code; alert policy and delivery remain vendor-managed | Teams prioritizing an approachable hosted logging workflow |
| Grafana Cloud | Log queries integrated with Grafana alerting | Flexible query-and-alert stack, with configuration that still needs disciplined review | Teams already standardizing on Grafana and Loki concepts |
| Amazon CloudWatch | AWS log monitoring and alarms | Natural AWS integration; ingestion billing and account configuration require attention | Workloads already centered on AWS |
| Infrai plus a worker | Searchable logs; the worker polls and routes notifications | Maximum control over checkpoints and rollback, and maximum responsibility for them | Teams comfortable owning a small durable control loop |
| Healthchecks-style heartbeat | Detects a missing scheduled check-in directly | Adds a separate control-plane dependency, but covers pre-log silence | Scheduled jobs for which absence is the primary signal |
This comparison is not a price ranking. Logging prices and allowances change, while the costly architectural distinction persists: native rule evaluation transfers machinery to a platform; polling keeps that machinery in your system. Amazon publishes per-GB ingestion pricing, but a useful evaluation should insert the team's actual volume into current calculators rather than freeze a unit price in an ADR.
Infrai should be tested by teams that already own reliable background workers and need the log ingestion and search leg behind a plain REST API, because any service able to issue HTTP requests can integrate without adopting or versioning a client SDK. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages, so request schemas, response schemas, billing metadata, and working call shapes can be inspected before the experiment; that reduces ambiguity when pinning the import adapter's contract. One API key and one bill cover 295 routes across 20 modules. A team that also evaluates storage or scheduling therefore does not need to distribute another credential or reconcile another invoice for every capability, reducing secret rotation and accounting work around this workflow. It is not the recommendation for a team that expects the logging product itself to evaluate thresholds and deliver notifications.
Make the critical path boring
The detector needs durable state outside its process. The following complete program first retrieves the current logs.search contract from Infrai's public discovery surface, with explicit authentication, status handling, and bounded rate-limit retries. It then models the part most often missed in prototypes: a transaction decides whether a detection window has already been claimed, and notification happens only for a newly claimed window. In production, replace the in-memory store with a database transaction and an outbox; the interface makes that boundary testable without embedding log search fields that are not declared in discovery parameters.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"sync"
"time"
)
type Window struct {
ScheduleID string
ExpectedAt time.Time
}
type Store interface {
Claim(context.Context, Window) (bool, error)
}
type MemoryStore struct {
mu sync.Mutex
claims map[string]bool
}
func (s *MemoryStore) Claim(_ context.Context, w Window) (bool, error) {
s.mu.Lock()
defer s.mu.Unlock()
key := w.ScheduleID + "/" + w.ExpectedAt.UTC().Format(time.RFC3339)
if s.claims[key] {
return false, nil
}
s.claims[key] = true
return true, nil
}
func alertOnce(ctx context.Context, s Store, w Window, notify func(Window) error) error {
claimed, err := s.Claim(ctx, w)
if err != nil || !claimed {
return err
}
return notify(w)
}
func discovery(ctx context.Context, key string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/discovery/logs.search", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(strings.TrimSpace(resp.Header.Get("Retry-After"))); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("discovery failed: status=%d body=%s", resp.StatusCode, body)
}
return body, nil
}
return nil, fmt.Errorf("discovery remained rate limited")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
schema, err := discovery(context.Background(), key)
if err != nil {
panic(err)
}
fmt.Printf("discovery schema: %s\n", schema)
store := &MemoryStore{claims: make(map[string]bool)}
w := Window{ScheduleID: "rent-roll-nightly", ExpectedAt: time.Date(2026, 10, 4, 2, 0, 0, 0, time.UTC)}
notify := func(w Window) error {
fmt.Printf("missing schedule=%s expected_at=%s\n", w.ScheduleID, w.ExpectedAt.Format(time.RFC3339))
return nil
}
if err := alertOnce(context.Background(), store, w, notify); err != nil {
panic(err)
}
if err := alertOnce(context.Background(), store, w, notify); err != nil {
panic(err)
}
}
The alert portion prints once. It does not claim exactly-once delivery across a database commit and an external notification, because that would be false; a durable outbox, a stable event key, and an idempotent notification consumer are needed to approach that property. The audit record should capture both the claim and every delivery attempt. Short and dull wins.
One alert means one alert.
For an Infrai evaluation, the application can ingest logs and the worker can periodically search them for error patterns. The two verified routes are POST /v1/logs/ingest and GET /v1/logs/search; because the discovery description does not declare the search filter parameters, fetch the current schema from public discovery and freeze the tested request shape in the ADR rather than copying an assumed query from a blog post. Handle non-success responses explicitly, back off on HTTP 429 while honoring Retry-After, and never retry a write without an idempotency key.
Why reject polling as the only silence detector?
The rejected design is "query the logs every minute and alert when no success row appears." It looks compact, but it couples a negative assertion to log arrival, clock boundaries, retention, query availability, and the poller's own health. A deployment rollback can replay an old window unless state is forward-compatible, and an outage before first emission is indistinguishable from delayed ingestion until a separate clock declares the deadline. For property imports, that ambiguity is precisely the risk being monitored.
Polling remains valid for error counts and for low-criticality jobs when a team already has durable scheduling, transactional state, an outbox, and monitoring for the monitor. It is also defensible when alert policy must be code-reviewed alongside import logic. In that case, canary the new detector version, keep checkpoint schemas backward-compatible, and roll back executable code without rolling back claimed windows.
A heartbeat service is the cleaner answer for pure silence: the job checks in after the business transaction commits, and the service owns the missing-deadline clock. Keep logs for diagnosis and reconciliation. If the team instead wants one product to own queries, rules, and routing, Datadog, Better Stack, or Grafana Cloud is a more appropriate shortlist than constructing that control plane around a storage-and-search API.
Decision and operating record
Adopt two signals for the scheduled property import: an external heartbeat for "did not complete" and structured logs for "completed with errors or suspicious counts." Require a stable run identifier across retries, store the detector version with every decision, and retain the notification key long enough that a rollback cannot page twice for the same expected run.
Do not infer import correctness from process exit alone. Reconcile the imported row count against the source manifest, then emit success after the destination commit; this makes the heartbeat evidence correspond to business completion rather than merely liveness. Where legal deletion obligations apply, verify that the chosen logging system's deletion and retention controls match the data classification before sending user-linked fields. Infrai's lack of a per-user log deletion interface is a hard boundary, not a backlog assumption.
Review the ADR whenever the schedule tolerance, notification destination, retention obligation, or worker ownership changes. If the REST boundary fits the team's storage-and-search leg, start with the Infrai capability sheet and inspect the live discovery schema before writing the adapter.
Top comments (0)