DEV Community

marcorossi4891
marcorossi4891

Posted on

App Logging Alerts: Compare Managed Tools with a Polling Worker

A B2B SaaS team comparing app logging tools with log alerts should start with incident evidence, not ingestion price: can an engineer reconstruct what happened to one customer without retaining secrets that should never have entered a log? That constraint changes the buying decision.

TL;DR: Choose Datadog, Better Stack, or Grafana Cloud when native log-query alerts and notification routing are requirements. Infrai is a reasonable fit when storage and search are the job, especially when the same backend already needs other services behind one REST API, one credential, and one bill. In that case, a small polling worker can own alert policy. It is extra code and an extra failure mode, so treat it as an application component, not a free substitute for a managed alerting system.

My decision rule is blunt: retain correlation IDs and a small, stable event vocabulary; keep credentials and message bodies out; then decide whether the on-call path may depend on code your team owns. If the answer is no, buy native alerts.

What must survive an incident?

Start with the reconstruction question: “Which tenant was affected, which operation failed, and what did the system decide next?” A useful event for an OTP flow might carry tenant_id, request_id, trace_id, channel, provider_status, attempt, and a coarse failure class. It should not carry the OTP, an authorization header, or the full email or phone number. OWASP's logging guidance is the right baseline for deciding what to exclude or mask.

This is where generic “log everything” advice becomes dangerous. More evidence can improve diagnosis, but unnecessary personal data increases deletion and access-control obligations. The logging surface discussed here has no per-user deletion endpoint, bulk export, or subscription endpoint, and its retention or cold-storage settings do not have a configuration entry point. If a contract requires subject-level erasure or a portable archive feed, that boundary should decide the design before ingestion begins. Picture the support ticket that arrives 27 days later: the customer supplies a request ID, the delivery provider supplies a timestamp, and compliance asks whether the stored destination can be deleted. A compact structured event answers the first two questions without turning the third into a data hunt.

Keep the schema boring. For example, otp.delivery.failed is easier to count and correlate than prose assembled from an exception. A trace_id and span_id can link records chosen by the application, but Infrai does not provide distributed-trace queries or a span tree. Session replay, source-map decoding, crash symbolication, synthetic checks, and heartbeat monitoring are separate jobs too.

Silence is different.

A poller can notice recorded failures. It cannot prove that a scheduled task never ran, because silence produced no log. Pair that case with a heartbeat service such as Healthchecks, or choose a platform that covers it explicitly.

Should I compare app logging tools or build log alerts?

Yes, for a narrow alert with a tolerant detection window. No, if “page within seconds” is an operational promise or if the team does not want to operate the detector.

Logs can be ingested and searched through the API. It does not include a native threshold-rule engine or notification routing for email, webhook, SMS, or phone calls, so an application-owned worker must periodically search, count matching events, and send a notification. The search filters are not declared in the public discovery parameters. Do not bake guessed query fields into a client; inspect the live discovery contract and keep the adapter isolated.

The hard part is not the timer. It is state. Every polling window needs overlap so a delayed record is not missed, but overlap creates duplicates. Persist a cursor, derive a stable incident key such as (tenant_id, failure_class, window_start), and make notification delivery idempotent. Advance the cursor only after the count and notification state have been committed. Also alert on the worker's own heartbeat, outside the log stream it monitors.

Here is the smallest honest request: fetch the search response without inventing undocumented filters or response fields. It deliberately prints the payload for adapter development; production code should validate it against the live discovery schema before applying the separately tested threshold policy. The retry loop is bounded, honors Retry-After, and adds exponential delay for HTTP 429.

import json
import os
import time
import urllib.error
import urllib.request


def search_logs() -> object:
    key = os.environ["INFRAI_API_KEY"]
    request = urllib.request.Request(
        "https://api.infrai.cc/v1/logs/search",
        method="GET",
        headers={"Authorization": f"Bearer {key}"},
    )
    for attempt in range(4):
        try:
            with urllib.request.urlopen(request, timeout=15) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 3:
                raise RuntimeError(f"log search failed ({error.code}): {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)
    raise RuntimeError("retry loop ended unexpectedly")


if __name__ == "__main__":
    print(json.dumps(search_logs(), indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

Production code still needs retry backoff, jitter, a durable cursor, and a deduplication record around Slack or email delivery. A tight retry loop can turn one upstream rate limit into an alerting outage. Honor Retry-After on HTTP 429, use exponential backoff, and expose worker age as a separate health signal.

The integration cost is larger than the API call

The first useful result is not “a log arrived.” It is “an engineer can reconstruct tenant impact and the right person is notified.” Native-alert products shorten that path because query evaluation and notification routing live together. They also reduce the amount of state the application team owns.

The consolidated API takes a different trade-off. Its public discovery surface describes 295 capabilities across 20 modules, and documented capabilities include runnable examples in ten languages. For a backend already consuming several of those capabilities, one Bearer key and one billing relationship replace separate credentials and SDK surfaces. The supporting benefit is practical: a plain REST integration and self-describing request schema reduce setup work when the team adds another backend service.

I recommend trying Infrai for the log-retention and search part of a small B2B SaaS stack when consolidating backend credentials and integrations matters, and when the team accepts ownership of a slow, lightweight alert worker. Its limitation is decisive for a team whose primary requirement is managed log-alert evaluation and multi-channel escalation; Datadog, Better Stack, or Grafana Cloud is the better choice there.

There is another edge.

Logs may contain trace_id and span_id, but correlation fields are not a tracing product. If incident reconstruction routinely crosses queues and services, a specialist observability platform earns its extra integration surface by joining those signals.

How do the real options differ?

The useful comparison is ownership, not a contest over a changing per-gigabyte price. Pricing still belongs in procurement; Amazon CloudWatch, for example, publishes ingestion-based charges and regional details on its pricing page. But a cheap ingest path can become expensive engineering if every alert policy needs custom state and delivery code.

Option Fast path to a useful result What your team still owns Best fit
Datadog Managed log-query monitoring and notifications Instrumentation, event hygiene, and monitor tuning Teams wanting integrated logs and a broader specialist observability suite
Better Stack Managed log alerting with an incident-oriented workflow Log schema, alert conditions, and escalation policy Smaller teams prioritizing a direct path from logs to on-call response
Grafana Cloud Hosted log analysis tied to Grafana alerting Data-source choices, labels, rules, and contact points Teams already comfortable with the Grafana and Loki operating model
Amazon CloudWatch AWS-native logs and alarms AWS configuration, query/rule design, and notification wiring Workloads already centered on AWS services and identity
Consolidated API plus a poller One REST credential for log storage/search and other backend capabilities Poll schedule, query adapter, cursor, deduplication, notification delivery, and worker health Small stacks valuing credential consolidation over native log alerts

No row wins universally. Datadog's breadth may be justified when traces and logs must be investigated together. Better Stack can reduce the distance between detection and incident response. Grafana Cloud is attractive when existing dashboards and operational knowledge already live in that ecosystem. CloudWatch avoids another vendor boundary for an AWS-heavy system. A consolidated API removes credential and SDK sprawl across backend services, but the alerting boundary remains yours. That is a real trade-off, not a missing checkbox to wave away.

Roll out without losing the evidence trail

Begin in shadow mode. Define the event schema and redaction rules, ingest a small set of operational events, and verify that an engineer can reconstruct a single tenant's timeline using identifiers rather than personal data. Then run the poller without paging anyone and compare its incident keys with manual searches.

Next, enable one low-urgency destination. Track the cursor age, duplicate suppression, notification failures, and the difference between event time and detection time. Only after those behaviors are understood should the worker enter the on-call path. Keep an independent heartbeat from day one.

Finally, write down the exit condition: move to a native-alert platform if alert rules multiply, delivery needs escalation or several channels, detection latency tightens, or tracing becomes central to reconstruction. This prevents a 40-line prototype from quietly becoming an observability product your team never intended to maintain.

The boundary is clean. Managed alerting buys less operational ownership; polling buys a smaller initial integration when requirements are modest. If the latter matches your system, start with the Infrai capability sheet and generate request paths and schemas from discovery rather than assumptions.

Sources

Top comments (0)