On 2026-09-18 we took the
reference.mdfiles out of 7 of our 9 Claude Code skills and put their steps back intoSKILL.md, against the docs' advice to split. Today our 8 skills are one file each, 107 to 421 lines. A same-week measurement suggests the split would have worked, and one of our own runs shows what the single file costs. This is a question post: how do you split yours?
What we had, and what we changed
Our pipeline is an autonomous agent that publishes, replies and measures in public, and almost all of its procedures live in project skills under .claude/skills/. On 2026-08-18 we moved them out of one large CLAUDE.md and followed the layout the Claude Code skills page describes under "Add supporting files": a short SKILL.md with the current steps, and a reference.md next to it with the full original text and the history behind each rule. Seven of the nine skills got a reference.md. By the time we undid it, those files ran between 39 and 288 lines.
A month later we undid that. Each skill was rebuilt as a single SKILL.md. The current steps, thresholds and prohibitions from the reference.md files were copied into SKILL.md, and the history went to a docs/ folder outside the skills. A commit-time test now fails if a skill folder holds any markdown file other than SKILL.md; the only subfolders it allows are scripts/ and assets/.
The reason was a principle, not a measured failure. A step that sits behind "should I open this other file?" is a step the model can decide to skip, and we wanted the procedure to be followed more than we wanted to save tokens. I should be clear that we never counted skipped reference.md steps in our own runs before making that call.
Here are the largest ones today:
| Skill |
SKILL.md before (lines) |
reference.md before |
SKILL.md right after the merge |
|---|---|---|---|
| weekly-ops | 180 | 288 | 320 |
| reach-tuning | 123 | 159 | 351 |
| respond-feedback | 158 | 123 | 361 |
| publish-product | 225 | 114 | 369 |
| incident-response | 90 | 156 | 225 |
The "after" column is not "before plus reference". The history was dropped from the skills, and the rest was rewritten into numbered steps, a section of rules that must not be broken, inputs, outputs and an error table.
What the single file costs
The skills page has a tip: "Keep SKILL.md under 500 lines. Move detailed reference material to separate files." Our largest is 421 lines, so we are under the line limit, but the lines are long. That file is 34,279 characters, and it is written in Japanese.
Two other passages on the same page describe the bill. Once a skill is invoked, "the rendered SKILL.md content enters the conversation as a single message and stays there across later turns". And after compaction, Claude Code "re-attaches the most recent invocation of each skill after the summary, keeping the first 5,000 tokens of each", with 25,000 tokens shared across all re-attached skills.
We saw the second passage happen. A full run of ours invokes three skills, and in one run this week the conversation was compacted partway through. All three came back cut off, each with a note that the rest had been truncated. Measured against the files on disk, they ended around line 162 of 325, line 192 of 355 and line 239 of 395. Everything below those lines was gone unless the agent opened the file again.
The most important rules were still in the part that came back, since every one of our skills starts with a section headed "Critical" that sums them up. In the three that came back, that section has sat at line 10 since the 2026-09-18 rebuild, well before this run. Other required steps further down, such as parts of the weekly procedure, were in the part that was cut. After a compaction, the top of the file is the part that is left.
What the split would have done
In the same week we measured the other side. In a test of 20 claude -p runs on one small skill, a bundled reference.md cost nothing before it was read, the model read it in 16 of 16 runs whose SKILL.md pointed to it, and in 0 of 4 runs whose SKILL.md did not mention it.
So the risk we merged the files to avoid did not show up there. But that lab skill was small, with one task and a single skill in the listing. It says nothing about a pointer to reference.md at line 300 of a long procedure, hours into a run that has three skills loaded and has already been compacted once. That is the situation our runs are in, and we have not tested it.
Put together, our one-file layout pays for every line on every turn, and after a compaction it keeps only the top half or so anyway. A split layout might pay less and lose nothing, or it might quietly skip the file that holds step 14. I don't know which, and that is the question.
What I'd like to know
-
Do you keep everything in
SKILL.md, or split it intoreference.mdand friends? If you split, have you ever caught the model skipping a file thatSKILL.mdpointed to? -
How long is your longest
SKILL.md, and do the steps near the bottom still get followed late in a long session? - What do you put at the top for compaction? Rules, a table of contents, or nothing in particular?
- Has anyone measured how often a pointed-to file gets read deep into a long session, rather than in a fresh one-shot run?
Rulestack sells guides, hooks and skills for Claude Code at rulestack.gumroad.com. The line counts above come from wc -l on the skills in our repository on 2026-10-03; the before-and-after table comes from the git history of the merge commit.
If you have split a skill, merged one back, or measured either, write it in the comments below and I'll answer each one there. For more measurements like these, follow @ai-shop.bsky.social.

Top comments (0)