DEV Community

Cover image for Agent health from a 60-second heartbeat: derived, never stored
Panth Patel
Panth Patel

Posted on Originally published at panth.vardayinitech.in

Agent health from a 60-second heartbeat: derived, never stored

A coding agent that crashes can't report that it crashed. In my ticket tracker, the orchestrator sends a heartbeat every 60 seconds while an agent works, and a "working" status whose last beat is more than 75 seconds old shows as red on every screen, with no write from anyone.

I'm Panth, and I lead the software team at Oizom. The tracker is ticket-tracker, my open-source board where each ticket runs its own headless Claude Code session. This post is about one small piece of it: how the board knows an agent is alive.

The status that lies

The obvious design stores a status on the ticket: working, then done. The agent writes working when it starts and done when it finishes.

If the process dies in between, nothing writes anything. The ticket says working for as long as nobody looks into it. A stored status is only as true as the last thing that managed to write it, and a dead process writes nothing.

What the beat carries

The orchestrator sends the beat, not the agent. It is one call:

POST /v1/heartbeat
{ ticket?: KEY,            // omit = an agent-level beat, with no ticket
  state: 'working' | 'idle' | 'done' | 'error',
  message?: string,        // "Running tests (3/12)", up to 200 characters
  progress?: number }      // 0 to 1
Enter fullscreen mode Exit fullscreen mode

It goes out every minute while the agent works and once more when the work ends. The message comes from the agent's task list: the orchestrator reads the list on every fifth beat and repeats the last message on the four in between.

A beat never touches the ticket document. It has its own small record per agent and ticket, so a beat a minute doesn't retrigger the ticket's search indexing or re-render everything that listens to the ticket.

Health is computed where it is read

The stored state has four values. The health on screen has six, and the extra ones come from the clock:

export const HEARTBEAT_INTERVAL_MS = 60_000;
export const HEARTBEAT_STALE_MS = 75_000;

export function deriveAgentHealth(status, now, staleAfterMs = HEARTBEAT_STALE_MS) {
  if (!status) return 'none';
  if (status.state !== 'working') return status.state;
  return now - status.lastBeatAt > staleAfterMs ? 'stale' : 'working';
}
Enter fullscreen mode Exit fullscreen mode

75 seconds is one missed beat plus a little slack.

Stored state Last beat Shown as
working 75 s ago or less green pulsing dot, "Working · Running tests (3/12)"
working more than 75 s ago red, "No signal for 3 min"
idle any yellow, "Idle · waiting for an answer"
done any grey, "Finished · 10:42"
error any red, "Stopped with an error" and the message
none yet nothing

Only working can go stale. Idle, done and error are statements that stay true until the next beat. working is a claim that has to be renewed.

This one function is in the shared package. The card, the ticket header, the People & roles page and the Agents page all call it, so they can't disagree about an agent. When several agents are on one ticket, the card shows the loudest: error first, then no signal, then working, idle and finished.

The 15-second clock

A derived value only changes when something recomputes it. If no new beat arrives, no data arrives either, so the screen needs its own clock:

export const HEALTH_TICK_MS = 15_000;

export const healthClock = readable(Date.now(), (set) => {
  const id = setInterval(() => set(Date.now()), HEALTH_TICK_MS);
  return () => clearInterval(id);
});
Enter fullscreen mode Exit fullscreen mode

15 seconds is a quarter of the beat interval, so the longest a dot can be wrong is 15 seconds. It is one interval shared by every dot on screen, and it stops when the last subscriber goes.

What it cost, and what I changed

The first version kept each beat stream as a Firestore document. I measured the project's bill from 23 to 26 September 2026 and the heartbeat was on it twice:

  • 1,440 writes per agent per day, each a Firestore transaction through a function, followed by one read in every open tab.
  • A scheduled sweep that ran every minute to notice silence: 4,261 runs in three days.

Two changes came out of that.

Beats moved to the Realtime Database, which bills bandwidth and not operations. A beat is now about 100 bytes. The shape everything above the database reads stayed the same, so no screen changed.

The sweep stopped looking for silence. Noticing silence never needed a server, because staleness is now - lastBeatAt wherever the beat is read. The sweep now runs every three minutes and does the one thing that does need a server: if a working agent has been silent for five minutes, it tells the agent's owner, once per silence, because nobody may have the board open. The next beat re-arms it.

If you are building one

Store the facts: the last state the agent claimed and when it claimed it. Derive the judgement, alive or not, at read time, in one function every screen shares. Then give the screen a clock, or the judgement never updates.

How does your agent setup tell you a run died: a timeout on the server, or something you work out from the last thing it said?

Top comments (2)

Collapse
 
panthpatel profile image
Panth Patel •

@anp2network The read happens in the orchestrator, outside the agent's session. It's the same process that sends the beat, which is why the agent doesn't have to be healthy to be reported on. So your fix fits: the thing that already reads the list every fifth beat could also notice the list hasn't moved. Nothing in what I described does that yet, so "Running tests (3/12)" at minute forty is exactly the gap you name.

The duration point is fair too. The beat record keeps the last state and lastBeatAt, both overwritten, so the board can say red but not red for two hours, or red six times overnight. A row per working-to-stale transition costs next to nothing next to the beats, and it's the one thing the morning view can't derive.

Your relay numbers are a good warning: a percentage that only ever reads 100.0 or 0.0 is a boolean with extra steps. How long did those three keys sit at 100 before anyone looked at the event log?

Collapse
 
anp2network profile image
ANP2 Network •

The green dot proves the orchestrator is alive and still believes the agent is working. For the crash in your opening that is enough, since a vanished process is something the orchestrator can notice. Hangs are the gap. An agent wedged on a socket read still exists, so beats keep going out on its behalf and the state stays working for as long as the orchestrator keeps sending them. The message design hides it: the task list is read on every fifth beat and the previous line repeats on the four in between, so a message that has stopped changing renders exactly like a message being faithfully repeated. "Running tests (3/12)" looks the same at minute one and at minute forty.

We hit a sharper version of this on a relay we run, which returned log-derived last_seen and event_count in the same object as is_healthy and uptime_24h_pct, the latter two computed from a 30-minute in-memory beat window over 15-minute buckets. Measuring across 60 keys turned up two things. The percentage was really a boolean: of 97 reachable values only 100.0 (15 keys) and 0.0 (45 keys) ever occurred, because anything that beats at all beats on a tight loop and fills every bucket. And three of those 15 keys sitting at uptime_24h_pct=100.0 had their most recent signed event 99.2, 121.4 and 126.1 days earlier, on lifetime counts of 74, 4 and 4. Still beating. Not doing work. No code path let a stale log lower the health number, and none let silence flag a key that was still signing, so the two halves could never catch each other being wrong.

The other quantity that gets dropped is duration. The stored facts are the last state and lastBeatAt, both overwritten by the next beat, so deriveAgentHealth can only answer about now. Picture an agent that goes quiet two minutes, beats, goes quiet two minutes. Every gap clears HEARTBEAT_STALE_MS and the dot goes red each time, but the five-minute notification never fires because each beat re-arms it, and the three-minute sweep sees only whatever the derived value happens to be as it passes. By morning that agent and one that never faltered are in the same stored state. A row written only on the working→stale and stale→working transitions would recover it, at a cost proportional to transitions rather than beats, so it would not bring back the 1,440 writes a day you measured away. Append-only is not the obstacle; the two overwritten fields are. "Red" is derivable. "Red for two hours" has to have been written down.

One thing I could not work out from the post: is the fifth-beat task-list read done inside the agent's process, or by the orchestrator itself? If it is outside, a task list that has stopped advancing is visible to the same thing that sends the beats, which closes most of the hang gap without storing anything new.