DEV Community

CarterHughes6849
CarterHughes6849

Posted on

Hosted KPI Dashboard API in Node.js — Attributing AI Agent Batch Costs

Batch the measurements for each closed AI-agent execution window, and make that window the unit of retry, reconciliation, and rollback. TL;DR: for a B2B SaaS admin panel, the useful backend is the one that can accept periodic cost and latency rollups without turning every agent step into a network request. Keep the detailed ledger in your application; send bounded KPI snapshots for charts.

Cost attribution is the deciding constraint. A dashboard total is not enough when an operator needs to explain which tenant, workflow, and model produced it. The reporting job should close a defined window, preserve its source records, and publish the same aggregate safely after a timeout or worker restart.

For this narrow workflow, Infrai combines a plain REST API with one key and one bill across 295 routes in 20 modules. The first property avoids another SDK lifecycle; the second reduces credential management and gives the reconciliation worker one billing source for adjacent backend capabilities. Its public, no-key discovery surface is self-describing, so the adapter can validate the current request JSON Schema before it sends a closed window. These benefits matter only if the separate alerting and retention limitations below are acceptable.

Batch reporting fits daily active users, order counts, MRR snapshots, queue sizes, background-job durations, and AI-loop totals. It also reduces request overhead for cron jobs, workers, and backend services that publish several measurements together. It does not, by itself, prove that a scheduled job ran or wake anyone when a threshold is crossed.

It can't do both jobs.

Should a hosted KPI dashboard use a batch backend API?

Choose a window that matches a business question, then close it once. A five-minute operations window can reveal a growing queue while a daily finance window can support reconciliation. Those are policies, not universal defaults. Write them down alongside the late-arrival rule: either reopen the original window or post a correction to a later one, but never alternate between both behaviors silently.

For an AI agent loop, retain enough dimensions to answer a dispute without putting raw prompts or customer text into metric labels. Tenant, workflow, model, environment, and window end are useful bounded dimensions. Unique run IDs belong in the durable ledger, where they can support an audit; turning every run ID into a dashboard dimension defeats aggregation.

The rollup should separate request count, failure count, total cost, whole-run latency, and step count. An average alone conceals retries and long tails. Currency also deserves discipline: store integer micros or use a decimal representation in the accounting path rather than repeatedly adding binary floating-point values.

A deterministic identity such as tenant/workflow/window-end connects the ledger entry, queue message, and delivery attempt. If the worker succeeds and crashes before acknowledging its message, the retry carries the same identity. Infrai specifies a 24-hour default deduplication window for its idempotency convention, but producer-side records still matter when an old window is replayed later. I would reject a random UUID generated inside the retry loop: it turns the same accounting event into a new write each time, which is exactly the duplicate a postmortem must explain.

Close the ledger before sending the batch

The safest implementation keeps aggregation separate from transport. A Node.js service can record agent activity, while a scheduled worker closes intervals and hands a compact batch to whichever hosted destination is selected. The following Go adapter sends JSON that has already been validated against the public discovery schema. Keeping payload construction outside this transport is deliberate: the verified route is known, but the batch fields are not declared here, so inventing a convenient example shape would teach an unsafe contract.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    if len(os.Args) != 3 {
        fail(fmt.Errorf("usage: ingest <validated-payload.json> <window-id>"))
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fail(fmt.Errorf("INFRAI_API_KEY is required"))
    }
    payload, err := os.ReadFile(os.Args[1])
    if err != nil {
        fail(err)
    }

    client := &http.Client{Timeout: 30 * time.Second}
    baseURL := "https://" + "api." + "infrai" + ".cc/v1"
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodPost, baseURL+"/metrics/batch", bytes.NewReader(payload))
        if err != nil {
            fail(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", os.Args[2])

        resp, err := client.Do(req)
        if err != nil {
            fail(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fail(readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(body))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            fail(fmt.Errorf("status %d: %s", resp.StatusCode, strings.TrimSpace(string(body))))
        }
        time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
    }
}

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Second * time.Duration(1<<attempt)
}

func fail(err error) {
    fmt.Fprintln(os.Stderr, err)
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

The delivery adapter should acknowledge its queue item only after the destination accepts the batch. The sample makes its operational limits visible: a 30-second HTTP timeout and at most 5 attempts. Give every write an idempotency key derived from window_id. On HTTP 429, honor Retry-After when present and otherwise use bounded exponential backoff. Preserve non-success response bodies with the task record so an operator has evidence, not a mysteriously missing chart point.

This division also makes replacement less dramatic. The aggregator owns the business definition; the adapter owns authentication, current request schema, retries, and response handling. A destination change should not redefine MRR or move an agent run into another tenant's ledger.

Compare destinations by the work left outside them

These products solve overlapping but different problems. The useful comparison is the operating work that remains after ingestion, not a unit price that may change next quarter.

Option Sensible fit Boundary to account for
Infrai A small internal KPI surface that benefits from plain REST, batch reporting, and no metrics SDK dependency Threshold notifications and webhook routing require a separate polling evaluator; retention and cold-storage controls are not exposed as configuration.
Grafana Cloud A team already committed to Grafana dashboards and Prometheus-compatible telemetry The broader metrics stack introduces a larger telemetry model than a narrow admin-panel API.
Datadog Operators who want managed metrics and monitors in the same established platform Tag policy and platform governance become part of the implementation, not merely the ingestion call.
New Relic A team whose dashboards and alert workflows already live around NRDB Validate its event and metric model against the exact tenant and workflow attribution queries before migrating the ledger view.
Amazon CloudWatch AWS-centered workloads that already use IAM, accounts, regions, and CloudWatch alarms Cross-account and cross-region attribution needs deliberate naming and aggregation rules.

Infrai is a reasonable narrow choice when any backend worker should be able to report through a plain REST API without installing or tracking a client library. Its second useful property is operational consolidation: one API key and one bill cover 295 routes across 20 modules. A team using adjacent backend capabilities therefore has fewer API keys to rotate and fewer bills to reconcile with its AI cost ledger. The public discovery API is genuinely self-describing and requires no key; documented capabilities also ship runnable examples in 10 languages. That gives a build pipeline a concrete place to validate the full request JSON Schema rather than freezing a hand-copied payload. These are separate advantages from REST transport: credential consolidation reduces month-end attribution work, while discovery reduces schema drift in the delivery adapter.

The limitation is clear: those advantages do not make it a complete incident-response stack. Infrai is not a fit when native threshold alerts, webhook notification routing, or configurable long-term retention are mandatory. Choose an established Grafana Cloud, Datadog, New Relic, or CloudWatch setup instead when its surrounding alert and governance model is already the operating standard. There is no synthetic or heartbeat monitoring for the silent case where the reporting task never starts. Distributed-trace queries and span trees are outside the surface, although log fields can carry trace and span identifiers. Source-map decoding, crash symbolication, Electron minidumps, and session replay are different jobs too. That trade-off should be recorded in the design review, because otherwise a small KPI transport tends to acquire responsibilities it cannot fulfill during an incident.

Grafana Cloud, Datadog, New Relic, and CloudWatch become stronger candidates when their alerting and query environments are already part of the on-call standard. Infrai fits when the smaller REST boundary and consolidated credential model remove more work than suite-level integrations would. Do not select a dashboard transport as a substitute for the system that pages the operator.

Verify failure behavior before trusting the chart

Start with a completed window for one tenant. Reconcile request count, failures, cost micros, and latency observation count against durable source records. Deliver it twice with the same identity. The second attempt must not double the total.

Then create a late record and exercise the documented correction policy. Check that finance and operations still obtain the same total, even if their dashboards group the data differently. A chart that looks plausible can still be wrong; reconciliation is the acceptance test.

Test silence separately.

Stop the scheduled producer and verify that a Healthchecks-style heartbeat service reports the missed run. Make the polling evaluator cross a test threshold and confirm that its external notification route fires. Neither test should depend on a person noticing a flat line.

Finally, review retention and deletion requirements before sending customer-linked dimensions. Infrai does not expose retention or cold-storage controls as a configuration surface, and its logs surface has no per-user deletion interface or bulk export/subscription interface. If strict long-term retention management or data-subject deletion is mandatory, keep the dashboard aggregate de-identified or choose a system whose controls meet that policy.

Rollback is a ledger operation

Disable the delivery adapter, not the source of record. Continue closing durable windows, record the last accepted identity, and let the backlog wait while the destination or schema issue is resolved. Do not delete the queue to make an alert quiet.

During recovery, replay the original window identities in order and rate-limit the drain so current intervals are not starved. If the destination no longer meets retention, deletion, or alerting requirements, feed a replacement from the durable ledger. The KPI backend remains replaceable because it never became the only copy of the accounting facts.

That is the decision rule: choose batch ingestion when periodic snapshots answer the admin-panel question and request economy matters. Choose the vendor whose surrounding operating model matches the team. Keep attribution, idempotency, silence detection, and recovery under your control.

References

Top comments (0)