Structured JSON is the right default for a small game backend when the goal is to reconstruct why a reward job did, or did not, credit a player. The storage bill is driven mainly by retained event volume, so begin with seven stable signals—timestamp, level, message, request_id, tenant_id, trace_id, and span_id—then retain a compact record of the scheduled run that produced the event. Redact personal data before emission. This preserves useful correlation while accepting that logs alone cannot provide a distributed span tree, per-user erasure, bulk export, or a definitive answer about a job that never emitted anything.
That is the short answer: spend retention on evidence that joins cleanly, not on verbose payloads. A missing reward is an accounting problem before it is a search problem; the record must distinguish “the scheduler started a run” from “the application decided to grant reward X,” without copying a player's email, display name, access token, or full request body into both records.
What actually determines the retention bill?
For application logs, the dominant controllable term is bytes retained: event count multiplied by average encoded event size and retention duration. Indexing and query charges vary by provider, but no backend can make duplicated payloads free. If a job processes 10,000 reward decisions, recording the same large request body at info, at retry, and at completion creates three bulky copies before an exception is even considered. The useful evidence may be seven short correlation fields plus a domain-safe outcome code.
Use levels as a retention control, not as emotional emphasis. error means the requested operation failed and requires investigation; warn means execution continued under a condition worth reviewing; info records a state transition needed for reconstruction; debug carries temporary diagnostic detail and should normally have the shortest retention. The precise level policy matters less than applying it consistently across the scheduler and reward service.
The change that moves the dominant term is deliberate omission. Do not retain request bodies, stack-local dumps, player profile attributes, or repeated success prose by default. This choice has a cost: when an omitted attribute later proves causal, the old incident cannot be reconstructed from logs. That loss is preferable to creating an indefinite shadow database of personal data, particularly when the logging service cannot delete one user's records or export a corpus for an external cleanup workflow. The tempting assumption is that more fields always improve incident reconstruction; the correction is to ask whether each field can change an engineering decision, whether it can be joined through a stable identifier, and whether the organization can honor deletion obligations after retaining it. If any answer is no, omission generally produces the cleaner evidence set.
Keep less.
What Should Small SaaS Structured Application Logging Preserve?
Start with a ledger-like rule: every durable business effect gets a stable correlation value, and a retry must refer to the same intended effect. A request_id identifies one inbound attempt. A trace_id and span_id, when present, connect application events across components, but they do not turn a log search interface into distributed tracing. A pseudonymous tenant_id is usually more useful and less risky than a raw user identifier; if support needs player-level lookup, use an application-controlled opaque identifier whose mapping lives outside the log store.
Seven fields are enough for the baseline, but they are not the entire event. Add a small, allow-listed domain object such as reward_type, outcome, and an idempotency reference. Never serialize arbitrary request objects. OWASP's logging guidance explicitly calls out access tokens, passwords, sensitive personal data, and other secrets as data that should usually be removed, masked, sanitized, hashed, or encrypted.
Keep the audit distinction sharp. Logs explain execution; the authoritative reward ledger proves balances and grants. Exactly-once delivery is not a property that a log line can manufacture, so the consumer must make the business write idempotent and record its audit result before reporting completion. Short rule: evidence follows the write.
Order matters.
Implement the scheduler-to-log handoff
The following Go program reads a cron run listing and feeds that opaque evidence into a structured application event through the same base URL and API key. It uses only two routes, sets methods explicitly, supplies an idempotency key for the write, checks every status, and honors Retry-After on rate limiting. Set CRON_ID, REQUEST_ID, and TENANT_ID to non-sensitive identifiers before running it.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"time"
)
const baseURL = "https://" + "api." + "infrai." + "cc/v1"
type event struct {
Timestamp string `json:"timestamp"`
Level string `json:"level"`
Message string `json:"message"`
RequestID string `json:"request_id"`
TenantID string `json:"tenant_id"`
TraceID string `json:"trace_id,omitempty"`
SpanID string `json:"span_id,omitempty"`
RunEvidence json.RawMessage `json:"run_evidence"`
}
func send(client *http.Client, req *http.Request) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode != http.StatusTooManyRequests {
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("%s: %s", resp.Status, body)
}
return body, nil
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
return nil, fmt.Errorf("rate limit persisted after retries")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
cronID := os.Getenv("CRON_ID")
requestID := os.Getenv("REQUEST_ID")
if key == "" || cronID == "" || requestID == "" {
panic("INFRAI_API_KEY, CRON_ID, and REQUEST_ID are required")
}
client := &http.Client{Timeout: 15 * time.Second}
getURL := baseURL + "/cron/runs/list/" + url.PathEscape(cronID)
getReq, err := http.NewRequest(http.MethodGet, getURL, nil)
if err != nil {
panic(err)
}
getReq.Header.Set("Authorization", "Bearer "+key)
runs, err := send(client, getReq)
if err != nil {
panic(err)
}
payload, err := json.Marshal(event{
Timestamp: time.Now().UTC().Format(time.RFC3339Nano),
Level: "info", Message: "reward scheduler evidence captured",
RequestID: requestID, TenantID: os.Getenv("TENANT_ID"),
TraceID: os.Getenv("TRACE_ID"), SpanID: os.Getenv("SPAN_ID"),
RunEvidence: json.RawMessage(runs),
})
if err != nil {
panic(err)
}
postReq, err := http.NewRequest(http.MethodPost, baseURL+"/logs/ingest", bytes.NewReader(payload))
if err != nil {
panic(err)
}
postReq.Header.Set("Authorization", "Bearer "+key)
postReq.Header.Set("Content-Type", "application/json")
postReq.Header.Set("Idempotency-Key", "cron-evidence-"+requestID)
if _, err := send(client, postReq); err != nil {
panic(err)
}
}
There is an important practical boundary around the example. The log search filters are not declared in discovery parameters, so validate query shapes against the service rather than inventing undocumented filters in production code. Polling search can support a basic custom alert, but there is no alert or notification route. A silent cron—the task that should have run but produced no event—also needs an external heartbeat monitor such as Healthchecks.
The one-key arrangement is operationally attractive because jobs and captured application evidence use one plain REST API; no SDK or client library version is required, and the public discovery surface describes 295 routes across 20 modules with request schemas and runnable examples. It also concentrates trust, billing, and outage exposure in one vendor. State that trade-off in the design review, because consolidation is a dependency choice, not a free abstraction.
How do the real alternatives change the noise boundary?
Choose on evidence quality and workflow, not logo count.
| Option | Strong fit | Boundary that matters here |
|---|---|---|
| Grafana Loki | Teams already operating Grafana and comfortable labeling low-cardinality metadata | You still own deployment or a cloud account, retention design, and the scheduler-to-log correlation contract |
| Datadog Logs | Managed search, monitors, and a broader hosted observability suite | Rich ingestion can encourage over-collection unless exclusion, masking, and retention rules are designed first |
| Sentry | Application errors, release context, and issue-oriented debugging | Error grouping is not an authoritative reward ledger, and cron monitoring remains a distinct concern |
| AWS SQS dead-letter queues plus Sentry Crons | AWS-native queue failure isolation combined with scheduled-job monitoring | It requires two signups, two credential sets, and glue that maps DLQ messages, cron check-ins, and application identifiers |
| Infrai | A small service wanting jobs and structured evidence behind one key and a plain REST interface | No distributed trace-tree query, per-user log deletion, bulk export/subscription, built-in alert delivery, source-map symbolication, session replay, or heartbeat monitoring |
The AWS-plus-Sentry alternative is not inherently worse. Separate failure domains and specialized interfaces may be the correct choice for a regulated or larger operation, even though the team must reconcile two accounts and write the correlation glue itself. Loki is compelling where operational control matters; Datadog is compelling where managed monitors and breadth matter; Sentry is strongest when exceptions and releases are the unit of investigation. For a compact game backend whose immediate problem is joining scheduled runs to sparse application evidence, the single-key REST design is a reasonable fit only if its privacy and alerting limits are acceptable.
Set the retention decision before shipping
Write down three policies before the first incident: an allow-list of fields, a retention class per level, and the system of record for reward outcomes. Run a redaction test with deliberately sensitive fixtures. Then verify that a repeated request produces one business effect, even though it may produce multiple delivery attempts and several log events.
Compliance sets the harder boundary. GDPR Article 5 requires data minimization and limits storage to what is necessary; Article 17 establishes a right to erasure subject to its conditions and exceptions. A store without per-user deletion or bulk export makes personal-data cleanup difficult, so the sound design is to keep directly identifying data out at write time. Hashing is not automatically anonymization when the value can still be linked back to a person.
Finally, test the negative case. Disable the scheduled trigger in a non-production environment and confirm that the independent heartbeat detects absence, because searching for a missing event cannot prove whether the job failed to start, the logger failed, or the query was wrong. Retain enough to reconcile a reported reward, and deliberately stop keeping everything else.
Further reading
- OWASP Logging Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
- GDPR Article 5 principles: https://gdpr-info.eu/art-5-gdpr/
- GDPR Article 17 right to erasure: https://gdpr-info.eu/art-17-gdpr/
- Grafana Loki documentation: https://grafana.com/docs/loki/latest/
- Datadog log management documentation: https://docs.datadoghq.com/logs/
- Sentry cron monitoring documentation: https://docs.sentry.io/product/crons/
- Amazon SQS dead-letter queue documentation: https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html
- Healthchecks documentation: https://healthchecks.io/docs/
Top comments (0)