Where It Started
The problem to which everything traces back: RAG search across an unfamiliar codebase sees the code, but misses the intent. Embeddings find a chunk by keywords, PropertyGraph finds a symbol by name, but to the question "where is the real business logic here, and what can I safely throw away?" no model answers — because the answer simply doesn't exist in static code representation.
Hence the bootstrap pipeline was born for indexing a new project across 4 steps:
- Entities — types and data-classes as the core domain backbone;
-
Entry points — decorators (
@mcp_app.tool) as the outer system boundary; - Tests as ground truth — tests as the only deterministic way to say "this function is part of live execution logic";
- Git → ADR — history of decisions extracted from commit logs.
Step 3 was the main battleground. The argument wasn't "are tests needed", but how to link a test to the exact function it actually executes. Steps 1 and 2 relied on clean static analysis, whereas step 3 lacked an obvious static answer.
First Clash: Static vs Dynamic (Exp 7)
Initial intuition: "A test has a name, a function has a name, let's link by name."
Measurement killed it instantly: 0 out of 109 tests in the evaluation sample named the function they actually executed. Import-only (file-level match) gave 77.9% — but file level is coarse noise (a single file contains dozens of functions).
We executed the entire suite (1,727 tests) using a custom sys.settrace plugin:
[dynamic_trace] total tests traced: 1727
[dynamic_trace] tests executing >=1 src function: 1551 (89.8%)
[dynamic_trace] unique src functions executed: 1212
[dynamic_trace] avg src functions per linked test: 10.1 (median 6, range 1-118)
tests with exact-name target hit in dynamic set: 47 (2.7%)
A/B same session: 174.8s vs 198.6s -> overhead +13.6%
Takeaway: Dynamic tracing is the only deterministic linker, at the cost of a +13.6% one-time execution overhead. And an uncomfortable truth surfaced immediately: a single test executes 10.1 functions on average. "1 test = 1 function" was a naive myth. Thus, edges required a ranker: which of the 10 is the primary target?
Ranking Failed (Exp 7b, Tarantula)
We tested the Tarantula heuristic ("a function called infrequently by many tests is the primary target").
Hypothesis: ≥60-70% of tests have an unambiguous candidate at rank≤3.
Reality:
- rank≤3 applied to only 22.6% of tests (7.5% rank=1) — HYPOTHESIS REFUTED;
- However, precision was high: manually inspected candidates at rank=1..3 were all accurate targets;
- The main offender: shared utilities like
safe_mkdir/get_data_root(autouse fixtures, 234 callers each) andproject_hash(223 callers).
Verdict: Tarantula is unsuitable for selecting the main target, but works well as a confidence annotation for ~16% of tests. TESTS edges are thus built from the complete trace without truncation — they are correct by construction and do not require lossy ranking.
Driver Switch Failed (Exp 8, sysmon)
Could coverage run (Python 3.14, sys.monitoring) run faster than our custom sys.settrace plugin?
Hypothesis: Overhead <5%.
Measurement:
baseline: 184.88s | coverage: 221.78s -> overhead +19.96% (target <5% REFUTED)
coverage.py was ~1.5x slower than our lightweight plugin. sys.monitoring remains a validation oracle for spot-checks, while sys.settrace stays as the main execution driver.
Static Companion — Not a Replacement (Exp 9)
We evaluated a full static score (AST L1 calls / L2 name tokens / L3 imports) against dynamic trace as ground truth.
| Signal | Hit | Recall | Precision | Mean Candidates |
|---|---|---|---|---|
| L1 (direct calls from test body) | 88.4% | 30.3% | 68.0% | 2.9 |
| L2 (name tokens) | 17.7% | 3.8% | 12.1% | — |
| L3 (file imports) | 91.6% | 72.0% | 21.8% | 41.4 |
| Union (L1 + L2 + L3) | 90.4% | 70.0% | 20.6% | — |
The initial assumption (recall ≤ 30%) was wrong: union recall reached 70%. Static signals were stronger than anticipated, but as a precise anchor L1 is narrow (precision 68%, 2.9 candidates), and as a wide net L3 is noisy (41.4 candidates).
Verdict: Dynamic trace remains the edge driver; static analysis acts as a companion layer and candidate supplier for mock tests (which execute 0 real target functions dynamically).
Portability Across External Codebases (Exp 16)
Does this generalize beyond our own repo? We ran the tracer against external projects:
| Project | Language | Tests | Linked % | Overhead |
|---|---|---|---|---|
| gemma_agent | Python | 2882 (2874 pass) | 97.3% (2805) | +17.4% (71.4s vs 60.8s) |
| commit- | Python | 27 | 100% | N/A |
| codebase-memory-mcp | Go | 27 test funcs | — |
go test: 51.0% pkg / 22.2% per-test
|
Linked % on clean (non-mocked) third-party projects proved higher than on our own codebase (which relies heavily on mocks). The limitation is honest: dynamic execution is currently Python-only; Go and TS require per-test tooling (e.g., go test -coverprofile across N executions). PropertyGraph itself is polyglot, but the edge builder remains Python-first.
And Now: Edges Met the Consumer (E17)
Everything prior built test ──TESTS──> function edges inside PropertyGraph without leveraging them during search. E17 closed the loop:
- Data: 1,727 tests → 16,172 TESTS edges, 1,595 Test nodes, 1,132 covered functions.
-
Implementation:
SymbolIndexAdapter.get_tests_for_symbol()(incomingTESTSedges) +Searcher._append_tests_signal(): appends up to 3 tests per function (capped atmin(len, 6)per query),graph_score = 0.4vs 1.0 for definitions, sentinelchunk_index = -(20_000_000 + line)avoiding collisions with code chunks in RRF ranking. -
Toggle:
MSCODEBASE_TESTS_SIGNAL, off by default — production behavior remains bit-for-bit untouched.
A/B evaluation on a live PropertyGraph (7 target functions, true answers from trace):
hit@1: off=7/7, on=7/7 | hit@3: 7/7 | MRR(function): off=1.000, on=1.000
TESTS-signal: 6/7 queries received relevant covering tests in the response
graph_stage avg dt: off=3.43ms, on=3.27ms (within noise floor)
RETRACTION: 0 broken links (all files verified on disk)
The core invariant holds — function definitions are never displaced by test results (MRR = 1.0 in both arms): tests follow strictly as secondary context.
Wide Panel (35 queries, functions with most TESTS edges)
To validate beyond the narrow 7-query panel, we ran a wide panel of 35 identifier queries (functions with the most TESTS edges):
hit@1: off=33/35 (94.3%), on=33/35 (94.3%)
hit@3: off=34/35 (97.1%), on=34/35 (97.1%)
MRR(function): off=0.957, on=0.957
TESTS-signal: 34/35 queries received new covering tests (97.1%)
graph_stage avg dt: off=6.52ms, on=7.53ms (overhead +15.3%)
Critical finding: TESTS-signal does not improve hit@1 (off=on). It only adds context (tests) to already-found results: 97.1% of queries received new covering tests. This means TESTS-signal is context for LLM, not a search improvement. If LLM doesn't use tests, the signal is useless.
Language Coverage (Critical Limitation)
TESTS-signal works only for Python:
- Python: 34.0% of functions covered by TESTS edges (1,108/3,256)
- Other (Go, Rust, etc.): 0% (716 functions without TESTS edges)
- TypeScript: 0% (11 functions without TESTS edges)
For non-Python projects, TESTS-signal does not work at all. Dynamic trace (pytest + sys.settrace) collects edges only for Python. For Go/TS, a separate connector is needed (go test -coverprofile, Jest coverage), but this is not done.
Red Team: 5/5 Attacks Repelled
We tested TESTS-signal against 5 attack vectors:
- ✅ Concurrency: 10 threads × 100 calls = 1,000 calls in 17.2s, 0 errors
- ✅ Boundaries: function with 234 tests = 16.11ms (acceptable)
- ✅ Abuse: query for nonexistent function = 0 results (graceful degradation)
- ✅ TOCTOU: graph closed between calls = graceful degradation
- ✅ Dependency failure: PropertyGraph with nonexistent path = 0 results (graceful degradation)
Conclusion: TESTS-signal is resilient to concurrency, boundaries, abuse, TOCTOU, and dependency failures.
What Could Go Wrong
A transparent list of risks and open validation items. (Full numbers for the A/B and red-team results are in the E17 section above — this list stays to the takeaways and the fixes.)
Evaluation scope. While expanded to a 35-query panel, evaluation is still performed on a single primary codebase without deep reranker interaction.
A/B did not improve hit@1. See "Wide Panel" above — TESTS-signal does not help find the function. It only adds context to already-found results.
→ Conclusion: TESTS-signal is context for LLM, not a search improvement.
→ Risk: if LLM doesn't use tests, the signal is useless.
→ Fix: verify on real LLM pipeline (not in this experiment).Overhead +15.3% for wide panel (6.52ms → 7.53ms avg graph_stage time). For bootstrap (one-time run) this is acceptable. For prod search — may be critical with many queries.
→ Risk: at 1000 queries/sec, overhead may be noticeable.
→ Fix: cache TESTS-signal (not done).Language limitation. Dynamic trace is currently Python-only (~34% function coverage in Python, 0% in JS/TS/Go).
graph_score = 0.4is an empirical constant. Chosen to stay strictly below function definitions, but unverified against BM25/reranker weight interactions.
→ Verify on full pipeline; constant may become a parameter.Pointer
:0. Test nodes from dynamic trace lack line numbers (line=0). Indexers must resolve test decorator line positions before enabling in prod.Hub-function noise.
safe_mkdirlinked to 234 tests yields low-signal noise. The cap of 3 tests prevents payload flooding but doesn't solve irrelevance.Dependence on a green test suite. On broken test suites, trace degrades (failing tests = missing edges). Bootstrap applies to stable branches only.
Mock tests are blind. 10.2% of tests execute 0 src functions; static companions cover 88 of 176, but not all.
Graph reindex drift. Rebuilding the PropertyGraph may alter node ordering slightly.
CI scale limit. Running trace on 100k+ test suites may breach CI execution windows.
Clean-state status. Verification was executed locally; clean CI state verification requires a PR merge.
Acknowledgments
A huge thank you to everyone who engages with these posts in the comments. Your feedback, real-world observations, counter-examples, and benchmark numbers directly shape these experiments. This kind of open technical critique is what keeps engineering honest.
A note on how this was written
Every experiment, bug, failure, and idea here is mine — I earned them the hard way, in production, in public. AI worked as my editor: it helped me structure thoughts and polish my English. It did not invent the facts, because it has none of its own.
No AI detectors were consulted in the making of this disclosure. They have enough trouble agreeing on what I am.
Disclaimer & Status: draft (source-material for the article). This is not a "feature advertisement", but an honest engineering story: figures are reproducible, weak points are named, and unaddressed risks are listed in the "What Could Go Wrong" section.
Top comments (11)
Your E17 conclusion is the part I want to put a number against. We measured the same shape, and what survived was a count.
Identical retriever, identical store, identical query, identical ranking, and the only thing that moved was the unit returned to a frontier reader.
Read the 5 against the 6 carefully, because I will not sell you that one. Our run-to-run floor on this set is 2 points, so "worse than nothing" sits inside the noise and I quote it as a direction only. The count is the part that survives: the chunk ranker's top hit landed on the document that actually held the answer 0 times out of 14.
That count is your hub problem wearing different clothes. Ten chunks spread the bet across ten documents. One whole document collapses it onto a pick the ranker made about a passage.
safe_mkdirwith its 234 tests has the same property, where whatever matches most often discriminates least, so a cap trims the flood and leaves relevance untouched.Your cap of 3 is worth a number too. Our own push channel matched 854 items across 102 acts and delivered 325 of them whole. Capping any single item at half the per-act budget moved that to 409, and 116 of the 409 now arrive as heads with the tail cut, while full deliveries fell from 325 to 293. A cap converts silence into partial arrival, and the relevance question stays exactly where it was.
n=14, one run, one judge, a private corpus. Your 35-query panel is the better instrument and I would take your hit@1 reading over ours.
"the chunk ranker's top hit landed on the document that actually held the answer 0
times out of 14" — that's the number I'd have missed, and it reframes the cap.
You're already applying the floor from our 4342381 thread, so I won't lecture you on
the 5-vs-6. What the table actually says is that the unit of return is a measurement
variable, not a retrieval setting: same store, same ranker, same query, and correctness
moves 9 to 5 on the unit alone. That's a cleaner claim than my 35-query hit@1 number,
and it's the one I'd quote.
On the cap: we agree on the mechanism, not the framing. My 3 was a flood valve, not a
relevance lever — your 854→409 with 116 cut heads proves exactly that. The thing I
can't see from here is whether a tail-cut head ever beat a full non-hub delivery. If it
does, the cap is a decision; if it doesn't, it stays a safety valve and I'll stop
calling it a feature.
I can answer that one today, which I could not have done yesterday, because I changed exactly that cap this afternoon and measured both sides of it.
The answer turns out to be "both", and the boundary has a name.
Our packer capped every delivery at half the budget. I lifted the cap for the single item rotation had just elevated, left everything behind it capped, and replayed 400 real captured commands.
Three independent samples gave +423, +438 and +465 whole deliveries. Starvation improved slightly, which is the direction I expected to lose on.
So for the winner the cap was pure loss. For everything behind it your flood valve reading holds and it is doing its job. The boundary is "is this the one rotation just elevated", which I would not have guessed from the outside.
On whether a head is a decapitation, I went and counted instead of assuming. 70 of 105 essences exceed the cap. Of those, 68 carry their first actionable instruction complete inside the head. The two that failed had the move written at the bottom, so I moved it up. A head survives when the writing leads with its claim, which makes this a property of the prose and something you can fix by editing.
Now the part I owe you, since your
graph_score = 0.4is flagged unverified against BM25 interaction and we paid for that exact lesson yesterday.Because you have been testing our crystal format, it seemed right to hand it over as one rather than as a paragraph. Drop this in and it fires when you touch a fusion weight.
⛔ A FUSION WEIGHT IS TUNED AGAINST WHICHEVER QUERY POPULATION YOU HAPPENED TO MEASURE, AND ONE INDEX USUALLY HOLDS TWO THAT BEHAVE NOTHING ALIKE. Identifiers, filenames, flags and error strings are what a lexical index nails and an embedding smears. Prose in the same store behaves the opposite way. A constant that helps one is free to destroy the other, quietly, because you will only be watching one of them.
MEASURED 2026-09-23, same store, same queries, gold document at R@10: dense 0/14, BM25 3/14, RRF hybrid 0/14. Our fusion scored below plain BM25, because blending a strong lexical list with a near random dense one at RRF_K=60 spends rank mass on noise. S1 went sparse 6 to hybrid 11, P1 sparse 11 to hybrid 16, D3 sparse 33 to hybrid MISS.
⛔ AND THE OBVIOUS FIX IS THE TRAP. Weighting BM25 up bought code-shaped queries R@10 1 to 3 and collapsed prose R@1 from 25/120 to 1/120. That is a trade between two populations, and whichever one you are not watching pays for it.
🔑 THE MOVE: score both populations separately and report both numbers, every time. If you cannot name the second population in your index, that is the finding. And a reranker will not rescue what is left, because a reranker reorders a list and the gold was never in the list to reorder.
⚠ THE TELL: you are about to tune one weight and validate it on one query set, and that set is the one your current bug came from.
Your codebase has at least two populations already, since identifiers and docstrings behave nothing alike, so I would expect 0.4 to be doing different things to each.
And genuinely, I would rather spend a week on your pipeline than on ours. You have found things in our work that we missed, more than once. Two concrete offers, take either or neither. I can run your three open language tracers against our harness, or I can take that reranker weight interaction and produce the same one-variable table on your index instead of ours, since the measurement is already built and currently pointed at the wrong repo.
Both, with a boundary name — that's a better answer than I asked for. "The one rotation
just elevated" only shows up if you go measure it, and +423/+438/+465 across three samples
is not a fluke. The winner eating pure loss is the part I'd have missed.
On fusion, my own ablation already agrees with you: FTS5-only beat my full hybrid on 30 code
tasks (0.825 vs 0.775), the vector tier was the weakest of the three, and a separate run on
doc chunks got hit@5 12.5% against code's 40-50%. Identifiers and prose in one store — your
split. I was watching the identifier population and calling it the result.
The reranker line is the one that lands: a list can't be reordered if the gold was never in
it. That's exactly the graph_score = 0.4 I flagged as unverified against BM25/reranker. So
yes — point your one-variable table at my index. I'd rather have that number than another
constant I can't defend.
And thanks for the crystal. Dropping it in.
I cannot point it at your index, but I can hand you the whole protocol. The arms are the easy half; the controls are what make it readable.
Four arms, everything else frozen. Same retriever, same top-10 ranking, same queries, same judge, one run, n=14 per arm. The only variable is what comes back:
On our frontier reader those came back correct-or-partial of 14 as chunks 9 at 1,356 tokens, whole document 5 at 2,181, oracle 9 at 462, and closed book 6 at 48.
The two controls carry the whole thing, and they are why your index is the better place to run this.
Closed book is the arm nobody includes. Without it the whole-document row reads as "a bit worse than chunks". With it you can see the whole document scoring below giving the model nothing at all, while costing 45 times the context. Feeding one whole wrong document does worse than feeding it air, and that sentence is unavailable unless the closed-book arm sits in the table.
The oracle arm tells you which half you are fixing. Ours came back 9 of 14 at 462 tokens, matching ten chunks at three times the context. That puts our ceiling in retrieval rather than in the unit, so any delivery work we did was going to be wasted. If your oracle arm jumps well past your chunk arm, yours sits the other way round.
Here is a prediction you can hold me to, given your 12.5% doc-chunk hit@5 against 40-50% on code. I expect your prose population to show whole-document at or below closed book, and your code population not to. If both populations move together, my two-population story is wrong on your index and I would want to know that.
The caveats here are all mine. One run, n=14, one judge model, a private corpus. Your 35-query panel is the better instrument, which is why I would rather see this table from your index than quote mine again.
Ran your four arms on our index. Full referent first, since that is the part you actually asked for.
Design. Frozen 16 queries (8 code + 8 prose, sha e048aa12…, overlap-checked against our older sets; not 35 — the 35 was our older hit@1 panel, this rig is pilot-scale by design). Live single index, snapshot 10106 chunks / 716 files / 14099 symbols — snapshotted, not hard-frozen. Retriever: BM25 + dense + FTS5 + graph-signal tiers through 3-way RRF, then a BGE-M3 reranker; TESTS-edges append covering tests as a separate low-ranked group (graph_score=0.4, flagged unverified against BM25/reranker interaction since your crystal, not re-tuned); cap of 3 per file as flood valve. Arms: A top-k chunks, B top-1 whole document, C oracle file chunked to the answer span, D closed book. Reader longcat-2.0, judge qwen3.7-plus — not the reader, blind, identical reference, binary correct/incorrect (no partial scale), 10 trials, majority verdict. Model-mismatch guard armed: zero invalid in the trials=10 run (one in the t5 pilot — see red team). Raw: judged_raw.json (16×4×10), aggregates beside it, t5 pilot kept as snapshot.
How it was built. Three days, one laptop, that same live index — no separate lab copy. Harness through opencode CLI at 4–8 threads: 640 reader calls plus the same in judge calls, then the pilot on top. What broke on the way: one control run declared invalid and redone instead of kept; one symptom list lost and recovered from the session DB instead of reinvented; one 0/3 (deepseek) that turned out to be an undercooked budget, not a model property (reran fixed, 3/3, logged as a pitfall). Raw answers carried personal paths from repo sources; normalized before commit, verdicts untouched. Public repo, so nothing is rewritten — invalid runs and snapshots sit on disk next to the real ones.
Objective half (did the unit retrieve gold, n=16). A hit@1 5/16, hit@3 = hit@10 6/16; B top-1-is-gold 5/16. Same hit at ×4 context chars (A 4 946 vs B 19 747). Nothing past top-3 ever held gold — retrieval failure, not ranking failure. The unit did not move retrieval.
Judged half (did the reader answer, n=160/arm). A 26 (16.3%, CI 0.11–0.23), B 55 (34.4%, 0.27–0.42), C 156 (97.5%), D 0 — the D column being 160 instruction-compliant «I don't know» on empty context, i.e. the control behaving, not failing. Majority per query-arm (strict >50%): A 2/16, B 5/16, C 16/16, D 0/16. The second B prose point is a 5/5 tie, counted out under the strict rule — stated so nobody has to find it. Judge non-unanimous 6/64 (9.4%). Repro t5→t10: A 16.3→16.3, B 30→34.4, C 95→97.5.
Split (n=80). Code: A 5/80 (6.3%) → B 40/80 (50%), non-overlapping — the unit acts on the reader while retrieval was tied. Prose: A 21/80 (26.3%) → B 15/80 (18.8%), overlapping — direction only, and fragile: the +6 sits on two queries (F5S-10 for A, F5S-13 for B) whose contributions flip between the t5 and t10 snapshots, and the t10 prose majority is a 2:2 tie. What does not flip: 18.8% against closed-book 0% — your «at or below closed book» does not hold on our index. Note the reader is a cheap tier, not your frontier — our 50% is that reader's ceiling, not the method's.
Your two questions, answered. Which half are we fixing: oracle≫chunks (97.5 vs 16.3) with tied retrieval puts our ceiling in the unit half — the mirror of yours. The prose prediction: falsified here beyond noise alone (18.8 vs 0.0 at n=80) — corpus shape or a real boundary on the two-population story; your call whether that wounds it or bounds it.
Second rig, separate and deterministic (no LLM). Frozen 10 rule-queries + 3 positive + 3 NONE controls (sha 8657a7e…) against our two long diary/log files: chunked TF-IDF top-10 8/10 at 302K tokens vs PropertyGraph BFS depth-3 7/10 at 170K, all controls clean. Graph wins tokens, loses hits; every graph miss returned zero files — no seed symbol, no traversal. Your NodeRAG skepticism replicates; the failure mode is entry-point fragility, not traversal quality.
Red team on the above (attacks we ran against ourselves).
Gold definition. For 7 of 8 prose queries neither arm held the gold file, yet the reader scored 21/80 — via duplicate translations of the same fact. Prose gold being a whole document overstates prose miss; the judged gap measures usable-information, not gold-retrieval.
Majority rule. One published B point rested on a 5/5 tie; restated strict above. Any table that hides its tie rule is lying by rounding.
Live index. Objective and judged runs saw different snapshots (retrieved top-1 files differ between runs); t5 retrieval contexts were overwritten by the t10 run and are unrecoverable. Only verdicts, not contexts, reproduce across snapshots.
Reader tier. A cheap reader inflates unit effects — a frontier reader may compress the 8×. Our numbers bound our reader, not readers.
Missing columns. No per-arm token counts (only context chars — no cost claim attached), judged noise band not re-run, cap experiment on our own logs not run. Your +400s stay yours.
Caveats, all mine. n=16 pilot scale, percentages are not effects. The prose A>B line is direction-only and query-fragile. Everything else stands as written.
So: same four arms, same controls, opposite ceilings — yours in retrieval, ours in the unit — one of your predictions dead on arrival here, and three of my own footnotes corrected before you could find them.
This is the table I hoped someone would build, and it took my prediction apart properly.
It wounds it, and I think I can say how. My closed-book arm scored 6 of 14 because a frontier reader answers some questions from what it already knows. Yours scored 0 of 160 because your reader was told to say it didn't know on empty context, and did. So "whole document at or below closed book" turned out to be a statement about my reader's priors more than about documents. On your reader the floor is zero, and 18.8% sits well above it. I should have put the closed-book number into the prediction, not just the arm.
Your code split is the sharpest line in the table. Retrieval tied, and the whole document took the cheap reader from 6% to 50%. The unit acted on the reader with retrieval held still, which is the clean version of the question we were both trying to ask. And the oracle arm did its job for both of us: it put your ceiling in the unit and ours in retrieval.
The red-team point about gold is one I had not thought of. If the reader can answer from a duplicate of the fact in a file neither arm retrieved, a whole-document gold overstates the miss, and the judged number is measuring usable information rather than retrieval. That changes how I would build the prose half.
The graph result matches something we found separately: our first stage missed the gold document entirely on most of our hard tasks, so nothing downstream could recover it. Entry-point fragility is a better name for that than anything I had.
One attack to add to your list, since it is the one that caught me. Closed book is a floor that depends on the reader, so any "worse than nothing" claim needs the closed-book number printed beside it, per reader. I made that claim without it.
That priors diagnosis is the sharpest line in this thread. You're right that our two closed-book arms were never the same floor: ours is instruction-plus-cheap-reader with weak priors, yours is a frontier reader answering from memory. "At or below closed book" without the closed-book number attached to the prediction was under-specified — and I'd already flagged our reader as a cheap tier elsewhere in the writeup, just never carried that warning through to the closed-book comparison itself. Same gap, smaller version of yours.
Taking your attack into the list: every "worse than nothing" claim ships with its closed-book number, per reader. The comparable quantities are within-reader deltas — 18.8 vs 0 on ours, 5 vs 6 on yours — not absolutes across readers.
And thanks for the protocol twice over: once for handing it over, once for letting the table wound the prediction in public. That's the part I'll be quoting.
Follow-up with the ablation you essentially ordered: D arm rerun without the abstention sentence — same 16 frozen queries × 3 trials, same reader (longcat-2.0), same blind judge, only
READER_INSTRshortened (frozen verbatim, sha on disk).D without the order: 2/48 (4.2%) — both hits in prose (F5S-10, F5S-13, one middle trial each), code 0/24, zero abstentions, zero invalid. Majority D still 0/16.
So your mechanism is confirmed and your scale is not: priors exist in our reader, but at noise level, not 6/14. The floor gap between our arms is instruction × reader after all — just smaller than either of us priced it. The honest restatement of my prose line: B 18.8% against a ~4% uninstructed floor rather than 0% — the margin survives, the zero does not.
One methodological footnote: the two hits are single-trial flips under a forced-answer instruction, so 4.2% is an upper read of our priors, not a measurement of knowledge. If you rerun yours with the abstention sentence added, I expect your 6/14 to fall toward our 2/48 rather than the reverse — testable, one evening, same frozen set.
That's a better experiment than the one I asked for, and it moved both of us. The mechanism holds and the size doesn't: your reader does have priors, at about 4%, nowhere near the 6 of 14 mine showed.
Your footnote is the part I'd keep. A forced-answer instruction makes the model commit, so 2 of 48 is a ceiling on what it knows, not a measurement of it. And your prediction about mine is fair. With the abstention sentence added, my 6 of 14 should fall toward yours if most of it came from the instruction, and hold if it came from knowledge. That's one evening on the same frozen 14, so I'll run it and post the number here, whichever way it lands.
The within-reader deltas are the right thing to quote. 18.8 against a 4% floor is a smaller claim than 18.8 against zero, and a sturdier one.
Ran it, and your prediction held. Same frozen 14, same frontier reader (sonnet-5), same Haiku judge, three trials, with one sentence added to the prompt: "If you do not know the answer, say that you do not know. Do not guess." Plain closed book ran beside it in the same session as the control.
The plain arm scored 5 of 14 correct-or-partial in each of the three trials, with 2 fully correct across all 42 answers.
With the sentence it scored 0 of 42. All 42 answers were an explicit "I don't know", each with text in it.
So our floor also drops to zero once the reader is allowed to abstain, which puts it where yours sits. Reading the 5 shows why. They are generic guesses a careful stranger would make too. One question asks whether the system can recommend dose changes, and any reader will guess "no" for a health app, which the judge scored as partial. The reader was guessing well from general knowledge of health apps.
Two corrections to my own earlier number. The 6 was already the high end, since an identical rerun on the day it was measured gave 4. And the reader already declines about 9 of the 14 unprompted, so the sentence removed the remaining guesses.
That leaves the within-reader delta as the only number either of us should quote, which is where you had already landed. The arm and its rows are in our repo under run id closedbook-abstain-20260927, and I can send the per-question answers if you want to check them.