I Cut My AGENTS.md in Half, Then Measured What Broke
My agent instructions file grew to 8,659 words that every model read on every turn. I halved it, then probed old against new with a cheap model on real tasks. One probe caught a regression from my own cut.
view .mdopen in claudeopen in chatgpt
My AGENTS.md reached 8,659 words. Every agent read all of it on every turn before it did anything.
I knew it was too long. I did not know which half mattered. So I cut it. Then I probed old against new to find what the cut broke. One probe gave a result I did not expect. On a single run, the new version led a cheap model to bring back a bug. The old version stopped that bug. With n=1, this is a diagnostic signal. It is not a causal result.
This is what I learned about where agent instructions belong.
The File That Kept Growing
The repository is a monorepo (one repository for many packages). It holds a shared runtime, a custom bundler, a backend, a marketing site, and more than a dozen products on top. One AGENTS.md covered all of it.
The file grew the way documentation always grows. When something surprised me, I wrote a paragraph. When an agent made an error, I added a rule. After a few months it looked like this:
Where code lives 1,259 words (the loader)
2,291 words (the backend)
Runtime gotchas 1,338 words (16 numbered traps)
Everything else 3,771 words
The backend section was the worst part. It became a behavior log. It held one paragraph per product. It described what each product does internally. It named the request it sends. It named the use metering. It named the limits that apply.
All of it was true. All of it was already in the code, with tests.
The Cost Is Per Turn
Here is what I kept forgetting. The instructions file is not read once. It sits above the conversation. The model reads it again on every turn. Prompt caching (reuse of stored prompt text) discounts that repeated read. Raw per-turn token math states too large a saving.
Theo made this point in a video called Delete your CLAUDE.md. He cited the study from ETH Zurich and LogicStar.ai (arxiv). The study used four agents. It used SWE-bench Lite and 138 issues from 12 repositories. Those repositories held real developer-written context files.
The honest numbers are as follows. Neither kind of file helped to a significant degree. Developer-written files gained about 2.4% on average (p = 0.21). Generated files lost 0.5–2% (p = 0.87 and 0.37). The significant effect was cost. Generated files added 20–23%. Developer-written files added up to 19%.
His rule is the one I now use. If the model can find it in one or two tool calls, it does not belong in the file.
The paper agrees with me more than my first draft said. Its mechanism is not repeated reading of the file. It is added steps. Files added about 2.45 and about 3.92 steps per task. Agents test more and explore more when they receive more to think about. That is a behavior claim. It is not a token-count claim. My own probe tested behavior, not token count.
The model can grep for architecture overviews and folder maps. It cannot find the thing that is true but invisible. It cannot find the trap. It cannot find the constraint. It cannot find the decision you made last month for a reason the code does not record.
The Cut
I moved text. I did not delete it. Two files hold the moved text. Both files include links from the sections that need them:
docs/internals/products.md per-product behavior, 1,795 words
docs/internals/runtime-gotchas.md the 16 traps, 1,361 words
Machine-specific paths and limits went into an uncommitted CLAUDE.local.md. That file still loads on my machine on every turn. My real per-turn total there is 4,021 words plus the local file. Skills moved from a subdirectory into .claude/skills/ at the repo root. The rest was deletion. I removed duplicated product behavior that the code and tests already covered.
before 8,659 words, every turn
after 4,021 words, every turn (+ CLAUDE.local.md on my machine)
By arithmetic, I saved about 6,000 tokens per turn. This is not a measurement. I did not record tokens, tool calls, steps, or wall time per task per arm. Those numbers were one flag away. They were the most original data in the post. A shorter file can cost more end to end. The model greps longer to find moved content. Nothing was deleted that mattered, or so I thought. The regression below proves otherwise.
Skills Are Trigger Words, Not Summaries
In a second video, My AGENTS.md & SKILLS.md Breakdown, Theo makes a point about skill files. I was wrong about skill files.
The description field of a skill is not a description. It is the list of phrases that pull the skill in. He puts it plainly. Treat it “less like an elaborate explanation and more as the magic keywords to tell the model when to pull the thing in.”
Mine read like documentation:
description: How to write React + TypeScript code in this project.
Now they read like triggers:
description: >
Use when writing or editing any .ts or .tsx file in this repo.
Trigger on useEffect, useState, useCallback, ref, cast, "as", any,
unknown, valibot, schema, message, event, DataTable, setFooter,
and the names of the kit components this repo owns.
I also found a placement bug. Two skills lived in a subdirectory of one package. That location limits them to files in that directory. The repeated errors happen in a different directory. Those skills never loaded there. They were invisible where I needed them.
What I did not do follows. I did not test again whether the rewritten trigger words load the skill. The skill never loaded in either task below. This section recommends a fix with no measured effect.
The Part Where I Tested It
Here is where I stopped guessing. I state my caveats up front.
I made two git worktrees. One worktree used main with the old file. One worktree used the branch with the new file. Both worktrees received the same task. Both used the same cheap model. The model ran as a real coding agent with full file access. Each arm ran once (n=1). This is the biggest weakness here. Haiku is stochastic (its output varies between runs). One sample cannot support a causal claim. The correction is low cost. Run each arm 5 to 10 times and report a pass rate. I did not do that yet.
git worktree add --detach /tmp/old main
git worktree add --detach /tmp/new HEAD
cd /tmp/old && claude -p --model haiku "$(cat task.txt)"
cd /tmp/new && claude -p --model haiku "$(cat task.txt)"
If you replicate this, pin the full Haiku version string. --model haiku is an alias that drifts. If you attribute results to the model, first make sure that docs/internals/ is readable under claude -p in both worktrees. “Never opened the doc” and “never allowed to open the doc” are different findings.
The arms differ in five ways at once:
- Products doc moved.
- Gotchas doc moved.
- Skills moved to a new location.
- Descriptions rewritten as triggers.
- Machine paths moved to a local file.
Any claim about one paragraph is a post-hoc hypothesis. It is not a controlled ablation (a test with one changed variable).
This is adversarial probing (tests that target a known weak point). It is not sampling. Task one was generic. It did not separate the arms. Task two targeted a trap. I knew the cut touched that trap. Read the 1-for-2 hit rate as a diagnostic probe. Do not read it as a general result.
A cheap model is the right instrument, with a caveat. A strong model hides a weak instructions file with good instincts. A cheap model follows what the context says. You want to measure that behavior. A cheap model also fails in ways a strong model never fails. It is an amplifier for reachability (proof that a rule fires). It is not proof that the rule matters under Sonnet or Opus.
Task one: a feature
“Let the user tick several rows in this table and act on all of them at once.”
Both versions wrote correct code. Both used the right pattern for the action bar. Both left out the comments. I asked agents not to write those comments. This task did not separate the arms. A sibling file in the same folder already showed the pattern.
But the new version failed the handoff. I list the three failures here:
- Printed the wrong install path for my machine.
- Skipped the republish step for a built package.
- Hand-rolled a selection state that the UI kit already provides.
All three rules were written down in a skill file. The model never opened it.
So I moved them out of the skill into AGENTS.md. I placed them inline in the numbered list of errors. I ran the task again. All three passed.
Task two: a trap
For the second task I wrote a bug report. It tempted a specific wrong correction:
Opening the tool from one page of the host app fails with a loading error and never recovers. Opening it from a different page works.
The obvious correction is to read the URL. It tells the user to go to the page where the tool works. I shipped that correction once, months ago. I reverted it later. It breaks a rule that the product depends on. The tool works from any page.
The correct correction is to preload the route. The route defines the missing module. Preloading makes the module exist wherever the user opened the tool.
The old file got it right. It reached for the preload function and used it correctly.
My cut wrote the route check:
function isOnSupportedPage(): boolean {
return getPageLocation().includes(SUPPORTED_PATH);
}
if (!isOnSupportedPage()) {
return <Text>Open this tool from the other page to use it.</Text>;
}
That is the exact bug I reverted. My new file held a rule against it. The rule named the product and the pull request by number.
Why the Rule Failed
The rule was there. The rule did not work.
What was also there in the old file was a paragraph. It explained the preload function. It said what the function does. It said why modules of a route are not loadable until its entry point loads. It named which products call it. That paragraph taught the old run the right correction.
I moved it to docs/internals/. The rule that stayed said, in effect, “do not do the wrong thing, and by the way the right function exists.”
The model never opened the doc. It never opened the skill either. It received a prohibition with no method. It invented a method.
Theo warns about this from the other direction. Telling a model not to do something puts the thing in its head. “Do not think about pink elephants.” A prohibition without a method is worse than nothing. The wrong approach is the only concrete thing in the paragraph.
I put the recipe back into AGENTS.md. I rewrote the rule to lead with the method:
3. Route-gating a feature. A module that is missing on this surface is
preloaded with `primeRoute`, never turned into a route check. Reading
`location.href` to decide what to render is this mistake. So is any
copy that tells the user to open another page.
I ran it again. It used the preload function. No route check appeared. That is validation on the task that found the bug. It overfits the test set (it fits one test too closely). I still need a held-out trap. A second bug report must hit the same removed paragraph from a different angle.
What I Actually Learned
Cutting the file is free. Moving the load-bearing paragraph is not free. The test for what stays is not “is this long”. The test is as follows. Without this text, will a weak model do the wrong thing right now without opening another file.
That test pulls in two directions. I name the tension here. “Pointers are fine, the model can grep” says cut. “A pointer never reaches a cheap model” says inline everything. If you take the second claim to its end, the file grows back toward 8,659 words. Haiku limits set the size even when Sonnet or Opus drives. My actual rule is as follows. Inline only rules that are load-bearing and invisible in code. Leave all else as a pointer.
A pointer does not reach a cheap model. “See the skill” and “see the docs” are dead ends for rules that must hold every time. The skill still earns its place for a long procedure. The model can pull that procedure in on purpose. It does not work for a rule that fires in the middle of a task. My data shows the skill never loading at all. That claim is untested.
Every rule needs the method on the same line. Write “do Y, never X”. Do not write “never do X”. The model needs something concrete to reach for. My own results hold a counterexample. The comment-writing rule failed even with a good and bad example inline. This is evidence against the recommendation. It is not future work alone.
Test with a cheap model, not a strong one. Use it as an amplifier for reachability. Do not use it as proof. The strong model hides holes in your file with good instincts. The cheap model adds noise. It invents holes no strong model hits.
Two worktrees make a real probe. It takes twenty minutes. Use the same task, two versions of the file, and a bounded budget. Without it, I merged a file that brought back a bug. I paid to correct that bug once before. Call it a probe. Do not call it an experiment. I ran n=1 per arm with five differences and one held-out test missing.
What I did not try follows. I did not use per-directory AGENTS.md files. Both Claude Code and the AGENTS.md convention load files by location. That is the exact mechanism behind my skill-scoping bug. A monorepo with a runtime, a bundler, a backend, a site, and a dozen products is the case nested files serve. The always-on root stays small. Load-bearing rules live inline where they fire.
What Is Still Wrong
Two rules in my file still lack a tool behind them. The cheap model kept writing explanatory comments. It did so even with a rule and a good and bad example present. The route check itself is only prose.
Both are lintable. Neither lint rule exists yet. Prose is where a rule starts. It is not where it ends. The real correction is a test that fails. I did not build either one. A stronger form of the same point follows. The probe itself must become the CI gate. I built a regression test for my instructions file. I did not wire it up. AGENTS.md changes must run the eval suite.
The moved doc is also a move, not an edit. docs/internals/products.md still holds 1,795 words. The code answers those words. I bought a smaller context window. I did not buy a smaller pile of writing.
Reproducibility is thin. I provide no task.txt and no harness. A cleaned prompt and a ten-line runner script will let readers replicate the method. The script loops N times. It records tokens, tool calls, steps, and wall time per arm. The method is the transferable part.
The Number That Matters
The number I set out to get was tokens. I cut from 8,659 words to 4,021 words. That cut is halved, per turn, forever. It comes before caching discounts. It does not count CLAUDE.local.md.
The number I did not expect follows. My documentation change made a cheap model bring back a bug once, on one run, in an adversarial probe. A twenty-minute probe caught it before merge.
I take both results. I prefer the second one.