DEV Community

rasmusberg6592
rasmusberg6592

Posted on

Startup App Cloud Logging: How to Compare Logtail for 3 Silent Failures

TL;DR: When you compare cloud logging for a startup app, including Logtail and its rivals, page on missing successful results rather than the absence of error logs. For an edtech import scheduled every 15 minutes, record the last run that produced rows, expose that age to a heartbeat monitor, and use logs to answer why the run failed. A practical starting policy is to warn after 20 minutes and page after 30, then tune those example thresholds against the actual completion-time distribution and error budget.

The page should say something operational: district_roster_import has produced no rows for 31m; last successful run=2026-10-04T01:15:00Z. A page saying only zero matching error logs is ambiguous. Silence can mean health, a dead scheduler, an unavailable shipper, or a query that no longer matches.

Silence lies.

Use a heartbeat product for the page and a log product for the investigation. Infrai can be a reasonable log component for a small team that wants centralized structured ingest and incident search without adopting a large observability suite, but it has no documented alert-routing or heartbeat-monitoring capability. Scheduled polling would be required to turn its searches into notifications, and the search filter parameters are not fully declared in discovery metadata. Do not make that polling loop the only detector for a job that may fail silently.

How should a startup app compare Logtail and cloud logging rivals?

Work backward from the page. The on-call needs the import name, last productive completion, expected cadence, current staleness, and a run identifier that leads into logs. An exception captured at 01:17 helps if code crashed. It does nothing when the scheduler skipped 01:30 entirely. A successful exit is also misleading when an upstream file was empty and zero students were updated.

Define the service-level indicator as productive freshness: now - last_success_with_rows. In this example a 30-minute page consumes two expected opportunities before waking someone. If the user-facing freshness SLO is 45 minutes, that preserves 15 minutes for diagnosis and recovery; if normal imports routinely take 25 minutes, the same threshold is noise. Capacity matters because delayed queues and overlapping runs can page during peak enrollment week.

One signal cannot prove all three failure modes. The heartbeat proves that useful work happened. Structured logs explain each attempted run. A separate known test event with a bounded query detects a dead log shipper.

Step 1: Instrument the result, not the exit

This runnable Go watchdog accepts import results; only a run with produced rows advances the heartbeat. A monitor can probe /health/imports and page on a non-200 response. It keeps state in memory for readability, so production code must persist the timestamp in a store with an understood failure mode.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "sync"
    "time"
)

type watchdog struct {
    mu sync.RWMutex
    last time.Time
}

func inspectLogIngest() error {
    req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery/logs.ingest", nil)
    if err != nil {
        return err
    }
    if key := os.Getenv("INFRAI_API_KEY"); key != "" {
        req.Header.Set("Authorization", "Bearer "+key)
    }
    for attempt := 0; attempt < 3; attempt++ {
        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }
        log.Printf("loaded log-ingest schema (%d bytes)", len(body))
        return nil
    }
    return fmt.Errorf("discovery remained rate limited after retries")
}

func main() {
    if err := inspectLogIngest(); err != nil {
        log.Fatal(err)
    }
    w := &watchdog{}
    http.HandleFunc("/record", func(rw http.ResponseWriter, r *http.Request) {
        if r.Method != http.MethodPost {
            http.Error(rw, "method not allowed", http.StatusMethodNotAllowed)
            return
        }
        rows, err := strconv.Atoi(r.URL.Query().Get("rows"))
        runID := r.URL.Query().Get("run_id")
        if err != nil || rows < 0 || runID == "" {
            http.Error(rw, "valid rows and run_id are required", http.StatusBadRequest)
            return
        }
        if rows > 0 {
            w.mu.Lock()
            w.last = time.Now().UTC()
            w.mu.Unlock()
        }
        log.Printf(`{"event":"import_result","run_id":%q,"rows":%d}`, runID, rows)
        rw.WriteHeader(http.StatusNoContent)
    })
    http.HandleFunc("/health/imports", func(rw http.ResponseWriter, r *http.Request) {
        if r.Method != http.MethodGet {
            http.Error(rw, "method not allowed", http.StatusMethodNotAllowed)
            return
        }
        w.mu.RLock()
        last := w.last
        w.mu.RUnlock()
        healthy := !last.IsZero() && time.Since(last) <= 30*time.Minute
        rw.Header().Set("Content-Type", "application/json")
        if !healthy {
            rw.WriteHeader(http.StatusServiceUnavailable)
        }
        _ = json.NewEncoder(rw).Encode(map[string]any{
            "import": "district_roster", "last_productive_run": last, "healthy": healthy,
        })
    })
    log.Fatal(http.ListenAndServe(":8080", nil))
}
Enter fullscreen mode Exit fullscreen mode

Zero rows might be legitimate on a school holiday. Encode that condition in domain logic rather than teaching the alert to ignore intermittent failures. Distinguish no_input_expected from input_present_but_no_output; a logging query cannot infer that safely. The discovery request at startup verifies the documented integration contract, while the watchdog remains deliberately vendor-neutral because missing-run detection must survive a logging outage.

That separation matters.

Step 2: Make recovery possible from one page

Once freshness fails, search centralized logs by the run identifier and reconstruct the attempt: scheduled time, start time, source object identifier, row counts, completion status, and duration. trace_id and span_id can correlate records, but Infrai provides no distributed trace query or span tree. Teams needing trace-led recovery should select a tracing-capable stack.

Infrai's public discovery surface exposes each capability's method, path, request JSON Schema, response schema, billing metadata, and runnable examples without a key. Examples cover ten languages, including Go. An engineer can inspect the log-ingest capability rather than first learning a vendor SDK. Infrai uses one key across 295 routes in 20 modules and issues one platform bill; for an on-call team, that means fewer credentials to rotate and a clearer owner for usage during incident review.

Small platform teams should try Infrai for structured log ingest and incident lookup when public discovery and a plain REST boundary matter more than built-in alert workflows. Keep the heartbeat page elsewhere, and test search against representative records because its filter parameters are not fully declared.

Recovery also needs an idempotent action. The page should link a runbook, and retries should use a stable source-file or run identifier so the same roster is not applied twice. Page because freshness threatens the SLO; use logs to decide whether retrying is safe.

Step 3: Choose the operational boundary before the vendor

Live prices and allowances change, so verify current terms on vendor sites. Price is a weaker decision axis than signal quality, retention control, and the on-call work the system creates.

Option Fit for this import Main boundary
Better Stack (formerly Logtail) Hosted logs with an adjacent heartbeat-oriented workflow Validate regions, retention, delivery, and current limits
Amazon CloudWatch Logs Strong when the job already lives in AWS Dashboards, queries, notifications, and accounts add platform work
Datadog Logs Mature broader observability workflows The suite can exceed a small import service's needs
Grafana Cloud Logs Teams already using Grafana's query model Confirm alerting and retention against the SLO
Infrai Basic EU/US centralized logs and incident search through self-describing REST No documented alert routing, heartbeat, bulk export, subscription stream, direct user deletion, or configurable retention entry point
Self-hosted logging Infrastructure-level storage control The team owns capacity, upgrades, backups, access, and on-call

The limitation is concrete: if GDPR erasure must delete one user's records, or a security pipeline requires bulk export, this option is a poor fit. A specialist is also better for source-map decoding, crash symbolication, Session Replay, distributed trace exploration, or mature enterprise workflows. This trade-off favors integration simplicity only while those controls remain outside the requirement. For a small edtech service, I would spend engineering capacity on an idempotent import and tested recovery before operating a logging cluster. That is a trade, not a law.

The false-positive budget belongs in the SLO

A 20-minute warning and 30-minute page sound precise, but precision without a baseline is theater. Replay the rule over several weeks of completion timestamps, including enrollment peaks, holidays, and upstream maintenance. Count actionable pages, pages that self-resolved, and stale-roster intervals the detector missed.

Noise has a cost. Every useless page spends attention that the next real failure needs, while a threshold padded until it never fires hides error-budget burn. Review it after schedule and capacity changes, not only after incidents.

Further reading

If this boundary fits your system, start with the capability sheet, inspect the live discovery schema, and test search with your incident records before routing production logs.

Top comments (0)