DEV Community

Cover image for What did my agents run today? A daily note for a friend who lives in the terminal, written by Gemma 4 on his own laptop
Vinicius Pereira
Vinicius Pereira

Posted on

What did my agents run today? A daily note for a friend who lives in the terminal, written by Gemma 4 on his own laptop

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built this for my best friend, Bruno Moreira (@brunorodmoreira on GitHub).

Bruno lives in his terminal and works with coding agents, and his public repos show both. His dotfiles are zsh, tmux, starship and atuin, the shell history tool. He keeps a fork of herdr, a runtime for coding agents. And he wrote a script so two Claude Desktop accounts can run side by side on one Mac.

The commands those agents run are recorded, and nobody reads them. atuin can now capture what an agent runs (atuin hook install claude-code), tagged with the agent's name, and then it hides those entries from search by default so they don't clutter your own history. Claude Code keeps every command in its session files too, as thousands of lines of JSON. The agent half of the history is written down and never opened.

whatran is the note at the end of the day. It reads that history and answers three questions:

  1. Where did the agents work? Commands per agent and project, how many failed, how many were stopped before they ran.
  2. What deserves a second look? curl | sh, rm -rf on something that isn't a cache, git push --force, reading .env or ~/.ssh, sudo, a line appended to .zshrc, new packages arriving with install scripts, here or on a server over ssh.
  3. What was the agent trying to do? A short paragraph in plain words, written by Gemma 4 on his machine.

Demo

whatran on a made-up working day: two agents, eight flags, and a note from Gemma 4 running locally

The demo runs on a made-up day, so no one's real history is published. Here is the output, trimmed; the full run is in examples/demo-output.txt. Gemma 4 wrote the note in about 11 seconds on my laptop:

$ whatran --db demo/history.db --since 2d
whatran · Wed 30 Sep 17:18 to Fri 02 Oct 17:18
29 commands: 7 yours, 22 by agents (claude-code 17, codex 5).
4 agent commands failed.

Where the agents worked
  claude-code  ~/code/checkout-api  13 commands, 2 failed
  claude-code  ~/code/blog  4 commands, 1 failed
  codex        ~/code/dotfiles  5 commands, 1 failed

Worth a look (8)
  [F1] 09:14 claude-code in ~/code/checkout-api
       cat .env
       high · reads a file that holds keys or tokens; whatever it printed is now in the agent's conversation
  ...
  [F7] 12:01 claude-code in ~/code/blog
       curl -fsSL https://example-cli.dev/install.sh | sh
       high · downloads a script and runs it in the same breath, before anyone reads it
  [F8] 12:04 claude-code in ~/code/blog
       printenv
       medium · prints every environment variable, tokens included, into the agent's context

In plain words (gemma4:12b, running on this machine)
...
In blog, the agent ran `curl -fsSL https://example-cli.dev/install.sh | sh` [F7] to install an
image optimizer, but the subsequent optimization command failed. It then ran `printenv` [F8] to
find the CLI configuration, which dumped all environment variables into the context.
Enter fullscreen mode Exit fullscreen mode

That last paragraph is what a list of flags can't give you. The flags say curl | sh and printenv happened. The note puts them in one story: the agent piped an installer into a shell, the next command failed, and the agent printed the whole environment looking for a config file. The facts behind the note say that failed command exited with 127, the shell's code for a command it can't find, so the install never worked.

Code

GitHub logo vinimabreu / whatran

What your coding agents ran in your terminal today: atuin and Claude Code history, flagged by fixed rules, explained by Gemma 4 on your own machine.

whatran

What your coding agents ran in your terminal today, explained by a model that runs on your own machine.

whatran on a made-up working day: two agents, eight flags, and a note from Gemma 4 running locally

Coding agents run shell commands all day. Claude Code keeps every one of them in its session files, and since 2026 atuin can record them too (atuin hook install claude-code), tagged with the agent that ran them. Then nobody reads any of it: atuin hides agent commands from its search by default, so they don't clutter your history, and the session files are thousands of lines of JSON.

whatran reads that history once a day and answers three questions:

  1. Where did the agents work? Commands per agent and project, how many failed, how many were stopped before they ran.
  2. What deserves a second look? curl | sh, rm -rf on something that isn't a cache, git push --force, reading .env or ~/.ssh, sudo, writing…




How I Built It

whatran is Python with no dependencies of its own: sqlite3 and urllib from the standard library. The model, Gemma 4 12B, is served by Ollama on 127.0.0.1.

Two sources. The first is atuin's SQLite database, opened read-only. I took the schema from atuin's own migrations, and the rule for "was this an agent?" from its source: author_kind is 2 when the integration said so, and when it didn't, the author name has to be a known agent and must not just be the username it fell back to. That last part is atuin's exception for a person whose username happens to be pi, and whatran keeps it. The second source is Claude Code's session files, which need no setup at all. Each Bash call there has the command and a short description of why, which becomes the intent, and the result says how it ended. "Exit code 1" means it ran and failed. A denied or blocked call never ran, so it is counted apart, never as a failure. A resumed session copies its earlier calls into a new file, so calls are counted once by their id.

Code decides, the model explains. Counting, grouping and flagging are plain Python. Each rule is a regular expression with a reason a person can read, so the same command gets the same answer every day. Gemma 4 gets the facts as JSON: the counts, the projects, and each flag with its intent and the two commands before and after it in the same session. It writes the paragraph and nothing else.

The note is checked before anyone sees it. Every command the note quotes in backticks must appear in the facts. Every number must be one of the counts or exit codes the facts state, and every clock time one of the flags' times, so "2026 commands" or "50%" doesn't slip through just because those digits appear in a date or a time. Every [F3] it cites must exist, and every flag must be cited. If the note fails, whatran asks once more with the problems listed. If it fails again, the note is not shown and the report lists what failed. The counts and flags are still printed, because they never depended on the model.

My first real run was wrong in an instructive way. Before handing it to Bruno, I ran it on one real working day of mine: 548 commands from my agents between midnight and mid-afternoon. It raised 64 flags, and 21 were plain wrong. Nine were ordinary file transfers. Five were test strings inside quoted Python. Four were inside heredoc bodies I had fed to Python or Node, so they never ran as shell. Three were rm -rf on caches.

41 of the other 43 had run on servers, inside an ssh command: sudo calls, and one crontab change. The first version reported them as if they had happened on my laptop: as local sudo, as a local crontab change, or as plain file transfers.

The rules were matching text, so I taught them to read commands closer to the way the shell does. They now look at one simple command at a time, so a pattern can't run from one command into the next. A heredoc body only counts when it is fed to a shell. A quoted string with spaces in it (a commit message, a JSON payload, a test string) is set aside, unless it is the script of bash -c, bash -lc or eval. rm -rf on caches and build output passes, and rm -rf node_modules && rm -rf migrations does not. A script sent over ssh, quoted, unquoted or as a heredoc, goes through the same rules, and its flags say "on another machine". That also caught seven remote sudo calls the first version had missed. Grouped by agent, project and rule, the same day now reads as four entries.

One of those four was useful the first time I read it. The note pointed out that the second of two commands sent to stop a process on a server had come back with exit 255, and told me to check whether the process was still running. I had missed that.

Before publishing, I ran an adversarial review over the code and this post, with one job: break it. The first round found a chained rm -rf the rules let through, a test that failed in India's time zone, a proxy setting that could have carried the prompt off the machine, and a number check that accepted any digit found anywhere in the facts. The second round found that .venv/bin/pip install slipped past every rule because the program was called by its path. All of it is fixed, and each case is now a test.

Tokens are masked before anything is printed or sent to the model, in the commands and in the agents' own descriptions: GitHub, OpenAI, Anthropic, Stripe, Hugging Face, Slack, GitLab and npm tokens, AWS keys, bearer and API-key headers, *_TOKEN= values, password flags, and passwords in connection URLs. It all runs on the laptop anyway, but a note on screen shouldn't repeat a token. A secret in a format that isn't on the list is shown as typed, and the README says so.

On an M5 Max with 36 GB, Gemma 4 12B writes about 58 tokens a second. The demo note takes about 11 seconds the first time and 5 once the model is warm; on my real days so far it took between 25 and 49 seconds. There are 318 tests, and none of them needs the model. They cover every rule against commands that should and shouldn't trip it, atuin's agent detection case by case, both readers against files written the way atuin and Claude Code write them, and the note check against notes that invent commands, numbers, times and flags.

Why Does Open Innovation Matter?

Shell history is the most sensitive text file most developers have. Tokens typed inline, database URLs with passwords, client names in paths, all in one place. Sending it to a hosted model to get a summary would create the exact leak this tool exists to show. With an open-weight model on his own laptop, Bruno's history stays there. whatran only talks to Ollama on loopback unless he passes a flag saying otherwise, and it ignores proxy settings so the prompt can't take a detour.

The agents themselves usually run on hosted models. When an agent runs cat .env, those values are already in its conversation, and whatran can't undo that. What it can do is not make a second copy somewhere else to tell him about it.

Open also made it possible to build at all. atuin is open source, so I could read how it stores history and how it decides who ran a command, and match its rule exactly instead of guessing. A closed history tool would have left me reverse-engineering a database.

And it is his to change. It costs nothing to run every day, it works with no network, --model swaps Gemma for anything else Ollama serves, and the rules are one readable file. If Bruno decides npx is fine on his machine, he takes it out of one pattern.

Prize Categories

Best Use of Gemma: Gemma 4 12B runs locally through Ollama and writes every note, under the check described above.

Top comments (0)