Claude Code, Codex and GitHub Copilot all keep a local record of what you did with them. If you bill by the hour, or just want to know what a feature really cost, that record is gold: every turn has a timestamp, and most have a model and token counts.
I've spent the last months turning those files into timesheets for Estela, an open-source CLI. The hard part wasn't the reporting. It was that each agent stores its history in a different shape, and each shape has a way to quietly give you the wrong number.
Here's what I found, agent by agent.
Claude Code: the same turn, up to five times
Sessions live in ~/.claude/projects/, one JSONL file per session. Each line is an event, and assistant turns carry a usage block with input, output and cache tokens.
The trap: a single assistant turn can appear in several rows (I've seen up to five) with identical usage. The file is a stream of updates, not a list of unique turns. Sum usage per row and your cost comes out about 2.9x too high on real sessions.
The fix is boring: dedupe by message.id before summing anything. Then, for time, group turns into work blocks: consecutive turns close together belong to the same stretch of work; a long gap starts a new one.
Codex: running totals and cache in the wrong place
Codex writes to ~/.codex/sessions/. Token usage arrives in token_count events, and each one carries two values: last_token_usage (this call) and total_token_usage (the session so far).
Trap one: sum the totals and you count every earlier call again on each event. On a real 68-event session that multiplied cost by 6. Summing the "last" values matched the final total exactly, which is a nice self-check to keep in your tests.
Trap two is subtler. Codex reports cached tokens inside input_tokens. Anthropic does the opposite and reports cache reads separately. If your code assumes one convention and reads the other, the cached part is billed at full input price and again at the cache price. On real data that turned $2.17 into $10.03.
GitHub Copilot: a log you have to replay
Copilot in VS Code keeps chats per workspace, under workspaceStorage//chatSessions. Older files are plain JSON. Newer ones are JSONL, and they don't contain the chat at all. They contain the changes that built it:
- kind 0: the initial state
- kind 1: set a value at a path
- kind 2: append to an array (optionally truncating first)
- kind 3: delete Read only the first line and you get an empty chat. You have to replay every change in order to reach the final state, and only then look at the requests. Once replayed, time is the easy part: each request has a start timestamp and a completedAt. In agent mode that span can be several minutes of real work that no timer would have caught. Two decisions I made for Copilot:
- Inline completions leave nothing on disk. That time can only come from commits, so it's labeled estimated, never measured.
- Cost stays empty. Copilot is a flat subscription. Multiplying tokens by API prices would produce a number nobody paid. An honest blank beats a precise-looking guess. ## The rule that ties it together Across all three, the principle I ended up with is: never mix measured and estimated without a label. Hours rebuilt from session timestamps are measured. Hours inferred from commits alone (hand-written code, or tools that leave no log) are estimates, and should err on the short side because someone is paying for them. Same with money. On a flat plan, the real AI cost of a project is your subscription split by how much each project used, not tokens at list price. ## If you're parsing these files yourself
- Claude Code: dedupe by message.id.
- Codex: sum last_token_usage, never the totals; subtract cached tokens from input.
- Copilot: replay the JSONL change log before reading anything.
- Write a test that compares your sum against the agent's own final total when it has one. Or skip all of it: Estela does this locally, no account, zero dependencies, MIT. npx estela setup Code: https://github.com/jdleonruiz/estela-cli Gemini CLI is the next agent I'd like to support, and there's an open issue for it. If you know how it stores sessions, a comment there would help a lot. And if you've hit other traps in these logs, I'd love to hear them.
Top comments (1)
the dedupe and the running-total traps are both in the parser. mine landed one step later, and your label rule is what would have caught it.
the counts were right when the script printed them. a second file quoted the first, and by the time I read the number back it was a copy of a measurement instead of one. it said 5. counting the source again gave 1, and nothing in the file marked which kind it was.
measured and estimated needed a third next to them for me. copied. a copy doesn't move when its source moves, and it reads exactly like the original.