DEV Community

Cover image for Everything that breaks once you stop watching the screen
Guillermo Leyendeker
Guillermo Leyendeker

Posted on Originally published at leyendeker.com

Everything that breaks once you stop watching the screen

An agent that works fine in a ten-minute demo and one that runs on its own for months are not the same system. The difference isn't architectural: it's dozens of edge cases that only surface with time on the clock.

Clocks drifting out of sync. A machine that suspends overnight. A parser that breaks on an unexpected prefix. A browser that grows without a ceiling until it eats the memory. A rhetorical question, written by the agent itself, that stalled the entire pipeline waiting for an answer nobody was going to give.

None of these appears in a diagram, and together they consumed as much time as every design decision before them. If you're about to leave something running unattended, this is the list you'd want to have read first.

Clocks

My favourite, because it's the hardest to suspect. For a while, the same task could read as alive or dead depending on which part of the system was looking at it.

The cause: the log file's name was generated in one timezone and its status was evaluated in another. Neither was "wrong" — they were simply different. The result was that a perfectly active run showed up as stopped, or the reverse, depending on the time of day.

Timezone mismatches are the kind of bug you don't find by reading code, because each piece looks reasonable on its own. You find it by watching the screen and noticing it says one thing while the system does another.

The machine that goes to sleep

The watchdog kills processes that have been running too long. Simple, until the machine suspends overnight.

On waking, the watchdog computed elapsed time in wall-clock hours and found processes with eight or ten hours of "activity" that hadn't executed a single cycle. And it killed them, correctly from its own point of view.

It had to be taught to subtract host suspension time from the calculation. It's a three-line adjustment that only occurs to you after losing work to it.

Fragile parsers

The end-to-end test progress bar sat frozen at zero. The suite ran fine, the tests passed, but the indicator never moved.

The reason: the test runner includes the project name on each output line when more than one project is configured, and my parser expected the format without that prefix. One extra token and the counter stayed at zero forever.

It's a reminder of something I badly underestimated early on: almost all integration with external tools runs through parsing their text output, and that output changes without notice, through configuration or version.

Memory that grows on its own

The browser running the interface reviews grew in memory without a ceiling. It wasn't a leak in my code: it was the normal behaviour of a browser held open for hours, accumulating tabs, contexts and cache.

The fix wasn't elegant and didn't need to be: recycle the containers when they cross a RAM threshold. Long-lived processes don't get fixed, they get restarted periodically.

States that lie on screen

There was an entire family of bugs around the interface showing things that weren't so.

Rows spinning indefinitely with an activity indicator, because the active task's identifier was compared incorrectly and no match ever switched it off. Live runs reading as stopped. Steps that did nothing but claimed the run button.

They all look cosmetic and none of them are, because the state on screen is what you use to decide whether to step in. An indicator that lies makes you wait for something already finished, or interrupt something that was working fine.

Operations that can't tell empty from broken

Read-only steps — the ones that explore code without modifying anything — still triggered the automatic push on completion. Since there was nothing to push, the operation failed, and that failure was reported as though the task had gone wrong.

It's the same pattern as the exhausted turns from the previous entry: the system couldn't distinguish "there was nothing to do" from "something went wrong". And that confusion generates noise which eventually teaches you to ignore the warnings, the worst possible outcome.

The rhetorical question that jammed everything

One skill ended its runs with a question along the lines of "shall we continue with the next part?". It wasn't expecting an answer: it was a closing formula.

But the orchestrator detects runs left waiting on user input, so they don't hang indefinitely. A question at the end, however rhetorical, triggers that detection and aborts the whole sequence.

The fix was to explicitly forbid closing on a question. It's also the case that gave rise to the self-improvement loop from the fourth entry: it was the canonical example of "I caught this by hand, the system should be catching it".

Other people's infrastructure falling over

When the provider's API returned a server error, the entire queue stopped. A thirty-second blip left the afternoon's work standing still.

Now those errors retry the step instead of killing the sequence. It's the same idea as in the entry on failure policy: you have to separate failures of the work from failures of the environment, because they call for different responses.

The pattern behind all of them

Looking at the full list, half of these cases weren't discovered by reading code. They were discovered by watching the screen and noticing something said one thing while the system did another.

And almost none of them are agent problems. They're old distributed-systems problems: desynchronised clocks, miscalculated timeouts, long-lived processes accumulating state, fragile parsers, external dependencies failing. They reappear intact because an agent running for hours is, for all practical purposes, an unreliable long-lived process.

If I had one piece of advice for someone about to build something like this, it would be: the interesting part — the prompts, the architecture, the retry policy — is the part you'll solve quickly. The part that eats your time is what keeps the system truthful when nobody is watching it.

The full arc

That's the path of these four months. It started as a screen with buttons for launching agents, with two commits devoted to one button's icon, and ended as a control system with distrust, budgets and auditing.

Today it has closed 2,095 fixes and features, plus 654 audits, on real projects.

The screen is still there. It's no longer the important part.

Top comments (0)