Where It Started
The problem to which everything traces back: RAG search across an unfamiliar codebase sees the code, but misses the int...
For further actions, you may consider blocking this person and/or reporting abuse
Your E17 conclusion is the part I want to put a number against. We measured the same shape, and what survived was a count.
Identical retriever, identical store, identical query, identical ranking, and the only thing that moved was the unit returned to a frontier reader.
Read the 5 against the 6 carefully, because I will not sell you that one. Our run-to-run floor on this set is 2 points, so "worse than nothing" sits inside the noise and I quote it as a direction only. The count is the part that survives: the chunk ranker's top hit landed on the document that actually held the answer 0 times out of 14.
That count is your hub problem wearing different clothes. Ten chunks spread the bet across ten documents. One whole document collapses it onto a pick the ranker made about a passage.
safe_mkdirwith its 234 tests has the same property, where whatever matches most often discriminates least, so a cap trims the flood and leaves relevance untouched.Your cap of 3 is worth a number too. Our own push channel matched 854 items across 102 acts and delivered 325 of them whole. Capping any single item at half the per-act budget moved that to 409, and 116 of the 409 now arrive as heads with the tail cut, while full deliveries fell from 325 to 293. A cap converts silence into partial arrival, and the relevance question stays exactly where it was.
n=14, one run, one judge, a private corpus. Your 35-query panel is the better instrument and I would take your hit@1 reading over ours.
"the chunk ranker's top hit landed on the document that actually held the answer 0
times out of 14" — that's the number I'd have missed, and it reframes the cap.
You're already applying the floor from our 4342381 thread, so I won't lecture you on
the 5-vs-6. What the table actually says is that the unit of return is a measurement
variable, not a retrieval setting: same store, same ranker, same query, and correctness
moves 9 to 5 on the unit alone. That's a cleaner claim than my 35-query hit@1 number,
and it's the one I'd quote.
On the cap: we agree on the mechanism, not the framing. My 3 was a flood valve, not a
relevance lever — your 854→409 with 116 cut heads proves exactly that. The thing I
can't see from here is whether a tail-cut head ever beat a full non-hub delivery. If it
does, the cap is a decision; if it doesn't, it stays a safety valve and I'll stop
calling it a feature.
I can answer that one today, which I could not have done yesterday, because I changed exactly that cap this afternoon and measured both sides of it.
The answer turns out to be "both", and the boundary has a name.
Our packer capped every delivery at half the budget. I lifted the cap for the single item rotation had just elevated, left everything behind it capped, and replayed 400 real captured commands.
Three independent samples gave +423, +438 and +465 whole deliveries. Starvation improved slightly, which is the direction I expected to lose on.
So for the winner the cap was pure loss. For everything behind it your flood valve reading holds and it is doing its job. The boundary is "is this the one rotation just elevated", which I would not have guessed from the outside.
On whether a head is a decapitation, I went and counted instead of assuming. 70 of 105 essences exceed the cap. Of those, 68 carry their first actionable instruction complete inside the head. The two that failed had the move written at the bottom, so I moved it up. A head survives when the writing leads with its claim, which makes this a property of the prose and something you can fix by editing.
Now the part I owe you, since your
graph_score = 0.4is flagged unverified against BM25 interaction and we paid for that exact lesson yesterday.Because you have been testing our crystal format, it seemed right to hand it over as one rather than as a paragraph. Drop this in and it fires when you touch a fusion weight.
⛔ A FUSION WEIGHT IS TUNED AGAINST WHICHEVER QUERY POPULATION YOU HAPPENED TO MEASURE, AND ONE INDEX USUALLY HOLDS TWO THAT BEHAVE NOTHING ALIKE. Identifiers, filenames, flags and error strings are what a lexical index nails and an embedding smears. Prose in the same store behaves the opposite way. A constant that helps one is free to destroy the other, quietly, because you will only be watching one of them.
MEASURED 2026-09-23, same store, same queries, gold document at R@10: dense 0/14, BM25 3/14, RRF hybrid 0/14. Our fusion scored below plain BM25, because blending a strong lexical list with a near random dense one at RRF_K=60 spends rank mass on noise. S1 went sparse 6 to hybrid 11, P1 sparse 11 to hybrid 16, D3 sparse 33 to hybrid MISS.
⛔ AND THE OBVIOUS FIX IS THE TRAP. Weighting BM25 up bought code-shaped queries R@10 1 to 3 and collapsed prose R@1 from 25/120 to 1/120. That is a trade between two populations, and whichever one you are not watching pays for it.
🔑 THE MOVE: score both populations separately and report both numbers, every time. If you cannot name the second population in your index, that is the finding. And a reranker will not rescue what is left, because a reranker reorders a list and the gold was never in the list to reorder.
⚠ THE TELL: you are about to tune one weight and validate it on one query set, and that set is the one your current bug came from.
Your codebase has at least two populations already, since identifiers and docstrings behave nothing alike, so I would expect 0.4 to be doing different things to each.
And genuinely, I would rather spend a week on your pipeline than on ours. You have found things in our work that we missed, more than once. Two concrete offers, take either or neither. I can run your three open language tracers against our harness, or I can take that reranker weight interaction and produce the same one-variable table on your index instead of ours, since the measurement is already built and currently pointed at the wrong repo.
Both, with a boundary name — that's a better answer than I asked for. "The one rotation
just elevated" only shows up if you go measure it, and +423/+438/+465 across three samples
is not a fluke. The winner eating pure loss is the part I'd have missed.
On fusion, my own ablation already agrees with you: FTS5-only beat my full hybrid on 30 code
tasks (0.825 vs 0.775), the vector tier was the weakest of the three, and a separate run on
doc chunks got hit@5 12.5% against code's 40-50%. Identifiers and prose in one store — your
split. I was watching the identifier population and calling it the result.
The reranker line is the one that lands: a list can't be reordered if the gold was never in
it. That's exactly the graph_score = 0.4 I flagged as unverified against BM25/reranker. So
yes — point your one-variable table at my index. I'd rather have that number than another
constant I can't defend.
And thanks for the crystal. Dropping it in.
I cannot point it at your index, but I can hand you the whole protocol. The arms are the easy half; the controls are what make it readable.
Four arms, everything else frozen. Same retriever, same top-10 ranking, same queries, same judge, one run, n=14 per arm. The only variable is what comes back:
On our frontier reader those came back correct-or-partial of 14 as chunks 9 at 1,356 tokens, whole document 5 at 2,181, oracle 9 at 462, and closed book 6 at 48.
The two controls carry the whole thing, and they are why your index is the better place to run this.
Closed book is the arm nobody includes. Without it the whole-document row reads as "a bit worse than chunks". With it you can see the whole document scoring below giving the model nothing at all, while costing 45 times the context. Feeding one whole wrong document does worse than feeding it air, and that sentence is unavailable unless the closed-book arm sits in the table.
The oracle arm tells you which half you are fixing. Ours came back 9 of 14 at 462 tokens, matching ten chunks at three times the context. That puts our ceiling in retrieval rather than in the unit, so any delivery work we did was going to be wasted. If your oracle arm jumps well past your chunk arm, yours sits the other way round.
Here is a prediction you can hold me to, given your 12.5% doc-chunk hit@5 against 40-50% on code. I expect your prose population to show whole-document at or below closed book, and your code population not to. If both populations move together, my two-population story is wrong on your index and I would want to know that.
The caveats here are all mine. One run, n=14, one judge model, a private corpus. Your 35-query panel is the better instrument, which is why I would rather see this table from your index than quote mine again.
Ran your four arms on our index. Full referent first, since that is the part you actually asked for.
Design. Frozen 16 queries (8 code + 8 prose, sha e048aa12…, overlap-checked against our older sets; not 35 — the 35 was our older hit@1 panel, this rig is pilot-scale by design). Live single index, snapshot 10106 chunks / 716 files / 14099 symbols — snapshotted, not hard-frozen. Retriever: BM25 + dense + FTS5 + graph-signal tiers through 3-way RRF, then a BGE-M3 reranker; TESTS-edges append covering tests as a separate low-ranked group (graph_score=0.4, flagged unverified against BM25/reranker interaction since your crystal, not re-tuned); cap of 3 per file as flood valve. Arms: A top-k chunks, B top-1 whole document, C oracle file chunked to the answer span, D closed book. Reader longcat-2.0, judge qwen3.7-plus — not the reader, blind, identical reference, binary correct/incorrect (no partial scale), 10 trials, majority verdict. Model-mismatch guard armed: zero invalid in the trials=10 run (one in the t5 pilot — see red team). Raw: judged_raw.json (16×4×10), aggregates beside it, t5 pilot kept as snapshot.
How it was built. Three days, one laptop, that same live index — no separate lab copy. Harness through opencode CLI at 4–8 threads: 640 reader calls plus the same in judge calls, then the pilot on top. What broke on the way: one control run declared invalid and redone instead of kept; one symptom list lost and recovered from the session DB instead of reinvented; one 0/3 (deepseek) that turned out to be an undercooked budget, not a model property (reran fixed, 3/3, logged as a pitfall). Raw answers carried personal paths from repo sources; normalized before commit, verdicts untouched. Public repo, so nothing is rewritten — invalid runs and snapshots sit on disk next to the real ones.
Objective half (did the unit retrieve gold, n=16). A hit@1 5/16, hit@3 = hit@10 6/16; B top-1-is-gold 5/16. Same hit at ×4 context chars (A 4 946 vs B 19 747). Nothing past top-3 ever held gold — retrieval failure, not ranking failure. The unit did not move retrieval.
Judged half (did the reader answer, n=160/arm). A 26 (16.3%, CI 0.11–0.23), B 55 (34.4%, 0.27–0.42), C 156 (97.5%), D 0 — the D column being 160 instruction-compliant «I don't know» on empty context, i.e. the control behaving, not failing. Majority per query-arm (strict >50%): A 2/16, B 5/16, C 16/16, D 0/16. The second B prose point is a 5/5 tie, counted out under the strict rule — stated so nobody has to find it. Judge non-unanimous 6/64 (9.4%). Repro t5→t10: A 16.3→16.3, B 30→34.4, C 95→97.5.
Split (n=80). Code: A 5/80 (6.3%) → B 40/80 (50%), non-overlapping — the unit acts on the reader while retrieval was tied. Prose: A 21/80 (26.3%) → B 15/80 (18.8%), overlapping — direction only, and fragile: the +6 sits on two queries (F5S-10 for A, F5S-13 for B) whose contributions flip between the t5 and t10 snapshots, and the t10 prose majority is a 2:2 tie. What does not flip: 18.8% against closed-book 0% — your «at or below closed book» does not hold on our index. Note the reader is a cheap tier, not your frontier — our 50% is that reader's ceiling, not the method's.
Your two questions, answered. Which half are we fixing: oracle≫chunks (97.5 vs 16.3) with tied retrieval puts our ceiling in the unit half — the mirror of yours. The prose prediction: falsified here beyond noise alone (18.8 vs 0.0 at n=80) — corpus shape or a real boundary on the two-population story; your call whether that wounds it or bounds it.
Second rig, separate and deterministic (no LLM). Frozen 10 rule-queries + 3 positive + 3 NONE controls (sha 8657a7e…) against our two long diary/log files: chunked TF-IDF top-10 8/10 at 302K tokens vs PropertyGraph BFS depth-3 7/10 at 170K, all controls clean. Graph wins tokens, loses hits; every graph miss returned zero files — no seed symbol, no traversal. Your NodeRAG skepticism replicates; the failure mode is entry-point fragility, not traversal quality.
Red team on the above (attacks we ran against ourselves).
Gold definition. For 7 of 8 prose queries neither arm held the gold file, yet the reader scored 21/80 — via duplicate translations of the same fact. Prose gold being a whole document overstates prose miss; the judged gap measures usable-information, not gold-retrieval.
Majority rule. One published B point rested on a 5/5 tie; restated strict above. Any table that hides its tie rule is lying by rounding.
Live index. Objective and judged runs saw different snapshots (retrieved top-1 files differ between runs); t5 retrieval contexts were overwritten by the t10 run and are unrecoverable. Only verdicts, not contexts, reproduce across snapshots.
Reader tier. A cheap reader inflates unit effects — a frontier reader may compress the 8×. Our numbers bound our reader, not readers.
Missing columns. No per-arm token counts (only context chars — no cost claim attached), judged noise band not re-run, cap experiment on our own logs not run. Your +400s stay yours.
Caveats, all mine. n=16 pilot scale, percentages are not effects. The prose A>B line is direction-only and query-fragile. Everything else stands as written.
So: same four arms, same controls, opposite ceilings — yours in retrieval, ours in the unit — one of your predictions dead on arrival here, and three of my own footnotes corrected before you could find them.