Two commands against the same data. Run them yourself:
$ curl -sS --get 'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production' \
--data-urlencode 'query=count(*)'
{"query":"count(*)","result":120,"syncTags":["s1:dmd/mg"],"ms":11}
120 documents, no key, no account. Now the path my agent actually uses:
$ curl -sS -o /dev/null -w '%{http_code}\n' \
'https://api.sanity.io/v1/context/organizations/<ORG>/mcp/self-correcting-systems'
401
Same underlying dataset. One route is publicly queryable. The route my agent takes goes through
an authenticated Context MCP endpoint, which needs an organization API token — not a project
token, which Sanity explicitly does not accept there. I had not provisioned Pouya any access to my
organization, so cloning the repo gave him no way to authenticate the path the agent actually uses.
I am being careful about "same" here: an MCP endpoint can be configured with its own sources and a
groqFilter, so it is not guaranteed to expose the same 120 documents the public route does. Same
source, different access path, different interface contract. That distinction turns out to be the
bug.
That gap is the whole story, and I did not know it was there until the contributor who fell into
it told me.
What happened
A developer named Pouya read a post of mine, decided my diagnosis was wrong, cloned the repo and
opened a pull request. First outside contribution the project has ever had.
PR #1 https://github.com/keniel13-ui/ask-the-record/pull/1
PouyaZX4:fix/groq-schema-routing
opened 2026-09-22T12:33:52Z
merged 2026-09-24T00:53:58Z -> 36.3 hours
3 commits, 1 file, +25 / -12, merge 4a2940f
It is merged, so it will not appear in the default pull request list, which shows open ones only.
A reviewer told me the repo had no PRs at all for exactly that reason, which is a small instance of
the same mistake this whole post is about.
His diagnosis was reasonable. My agent queries a Sanity dataset with GROQ, and he thought the
failures came from the model guessing at document types and inventing query syntax. His fix: put
the real schema and worked examples into the tool description, so the model is told the shape
instead of guessing it.
Sound idea. I would have tried the same thing.
Then I ran his examples.
The two queries
Both from the pull request description. Here is the one for a claim, against the live dataset:
*[_type == "claim" && _id == "claim-ledger-population"][0].{status, expiryStatus}
-> HTTP 400 "attribute expected"
A stray dot between [0] and {. Remove it and it works:
*[_type == "claim" && _id == "claim-ledger-population"][0]{status, expiryStatus}
-> {"status": "standing", "expiryStatus": "no_expiry_set"}
And the one for a patch record:
*[_type == "patch" && findingRef._ref == "finding-b1"][0]
-> null
Null for two separate reasons. There is no findingRef field on my patch documents — it is
findings, and it is an array. And the ids are case sensitive: finding-B1, not finding-b1.
Written correctly:
*[_type == "patch" && "finding-B1" in findings[]._ref][0]{sha, inMain}
-> {"sha": "dd1a654", "inMain": false}
Six field names that do not exist
The patch did not just carry two broken examples. It wrote a schema into the tool description,
and the schema did not match my data:
| the patch told the model | what the database actually has |
|---|---|
claim.subject |
nothing |
claim.statement |
text |
finding.finders |
foundBy |
finding.verifiedReceipts |
commentId or commentIds, depending on the record |
patch.findingRef |
findings[] |
patch.commitHash |
sha |
Check it yourself, no key required:
curl -sS --get 'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production' \
--data-urlencode 'query=*[_id=="finding-B1"][0]'
That is the record his example referenced. An earlier draft of this post ran the same command
against finding-B8 instead, because B8 carries commentIds and B1 carries the singular
commentId, which made my table read cleaner. That is picking the sample that proves the point,
in a post about not doing that, and a reviewer caught it before it shipped.
So a patch written to stop a model from inventing schema would have installed a mismatched schema
into the exact place the model reads it from. Worse than the bug it fixed, because a guess made at
runtime is a guess, while a guess sitting in the tool description arrives with the authority of
documentation.
Why the names were wrong, which is not what I assumed
I wrote a draft of this post that said I did not know how the patch was produced and was not going
to guess. That was the right call, because when I asked, the answer was nothing I would have
guessed. In his words, published with his permission:
The dataset itself is public of course, but I wasn't able to replicate the full MCP setup on my
end at the time I tested it so I was hitting that 401 on the Context MCP endpoint. So I just
tested the query generation logic against an offline mock schema on LM Studio instead, which is
where those field names came from.
Nothing was hallucinated. He hit the 401 at the top of this post.
He could read every one of my 120 documents in a browser. He could not run my agent, because the
agent does not take the public route — it goes through an org-scoped MCP endpoint needing an
organization credential I had not provided and was not going to publish with the repository.
So he did the careful thing. He built an offline mock of the schema and tested the query
generation logic against that. And when he built the mock he chose better names than mine:
When creating that schema, I used clean, self-describing semantic names (commitHash instead of
sha, verifiedReceipts instead of commentIds, finders instead of foundBy) because explicit naming
makes it way easier for the SLMs to understand what fields represent and prevents confusion.
He is right about that, by the way. commitHash is a better name than sha. foundBy is worse
than finders. Every wrong name in that table is the name a competent person would pick if they
were designing the schema rather than reading it.
One of them is worse than a naming difference. verifiedReceipts does not just rename
commentIds — it asserts something the data cannot support. A comment id says a comment exists at
that address. It says nothing about whether anyone verified it. That distinction is the entire
subject of the record it was describing.
The defect is mine
The mock is not the bug. The mock was a reasonable response to a 401.
The bug is that my repository advertised a reproduction path that is not the one the system
uses. The README says "Public dataset" and gives the URL at the top of this post. That is true,
and it is what I put in an earlier version of this very post — a curl, no key, check it
yourself. It reproduces the data. It does not reproduce the agent, which reaches the same
records through an endpoint that returns 401 to everyone who is not me.
A contributor following my README lands in a place where every record is visible and nothing is
runnable. The only way forward is to build a stand-in. And a stand-in built at that boundary does
not stay at the boundary — his mock's field names travelled from a local LM Studio test into the
tool description, which is where my harness explicitly tells the model what the schema is.
That is the whole difference in one line. A runtime guess is visibly a guess. Put the same guess in
the tool description and I have promoted it into authoritative guidance.
I have shipped the same class of error. I described my own validator to someone as checking
whether a cited record matched what was retrieved. It does not. I was describing the system I
meant to build, from memory, instead of opening the file.
What actually resolved it
Not an argument. Two commands.
I pulled his branch, ran both examples against the live dataset, and sent him exactly what came
back: the 400, the null, and the corrected queries with their real output. No opinion about his
approach, no debate about the diagnosis.
He pushed a revision in about two hours. Every field name corrected, both queries rewritten:
claim claim-ledger-population -> standing / no_expiry_set
patch for finding-B1 -> dd1a654 / false
finding finding-B8 -> commentIds ['3ee98'], status unbuilt
All three still return exactly that today. Note the third is B8, not the B1 I told you to curl
above — B1 carries the singular commentId "3eanf" and status implemented. Those are two
different records and I have mixed them up once already while writing this.
Still one stale reference in a routing rule, so I sent that too, with the null it produced. He
fixed it that morning. I re-ran the three patterns, checked the file still parsed and still had no
third-party imports, and merged it.
Three commits, thirty six hours, between two people who have never met. And one detail I like:
the pull request body still shows [0].{status, expiryStatus}, the broken stray-dot form, after
the merged code had moved to the corrected one. Documentation can keep a false schema alive after
the executable stops using it, which is the same failure as this whole post, one layer up.
The thing worth taking
If you maintain anything an outsider might contribute to, run this check:
Can someone who clones your repo actually execute the path your system takes, or only the path
your README documents?
For me those were different, and the gap was invisible from the inside because I hold the token.
Everything worked on my machine for a reason I never had to think about.
When they differ, a contributor's only option is a mock. They will build a good one — Pouya's
names were better than mine. And then the mock's assumptions become your documentation, because
the tool description is documentation, and the model does not know it was written against a
fixture.
There is a third option, and it is better than either of the two I first wrote down. I said the
fix was to make the real path reachable, or to say in the README that it is not. Both are weak,
because both still leave the contributor inventing the contract.
If outsiders cannot execute a privileged dependency, ship them a reproducible contract for it.
A fixture generated from the real schema, checked into the repo, lets someone test against the same
interface without ever receiving my organization credential. The pipeline becomes
production schema -> generated contract fixture -> contributor harness
instead of
README -> unreachable MCP -> contributor invents a substitute schema
The problem was never that Pouya mocked the boundary. It is that my repository gave him no
canonical boundary to mock, so he had to design one — and a well-designed guess is still a guess.
That generalises past Sanity and past agents. Anything an outsider cannot run — a private API, a
payment sandbox, an internal queue, an OAuth service — has this shape. If you do not own the
stand-in, your contributors will build one, and theirs will encode their assumptions instead of
yours.
Thanks to Pouya for the patch, for taking two rounds of corrections without once arguing the
diagnosis, and for answering the question about where those names came from when he could easily
have let me publish a guess instead. The merge is 4a2940f. Every query above can be run by
anyone against the public dataset — which, as it turns out, is exactly the point.
Top comments (12)
The mock becoming documentation is a useful signal: the missing artifact was not another prose guide, but an executable example of the tool contract. I would keep the successful path, one realistic failure path and the expected postconditions together. Then an agent can test whether it selected the right endpoint, sent an accepted shape and produced the intended external state. A request returning 200 is weaker evidence than a read-back showing that the right object changed.
Agreed on keeping the success path, one realistic failure and the expected end state together. In this case they were all read queries, so there wasn't a changed object to read back. What settled it was checking what actually came back against what each query said it would fetch. The 400, the null, then the corrected ones with real values. A 200 by itself wouldn't have told me anything there either.
That makes sense. For reads, the recovery check is semantic rather than stateful.
The contract could be: query intent → expected shape and content → actual returned values, while keeping 400, null and plausible-but-wrong results as separate outcomes. I’d also preserve one negative control: a query that should return nothing. Otherwise “real values came back” can pass even when the filter or tenant boundary is wrong.
the generated fixture is the useful fix here. i'd make its schema revision visible and fail the contributor tests when it no longer matches the published contract. otherwise today's canonical fixture becomes tomorrow's very convincing old mock. the singular commentId versus commentIds is a good regression case because a neat rename would quietly erase a real difference.
The version stamp is the piece I was missing. A generated fixture with no schema revision on it just turns into the next really convincing old mock. Failing contributor tests when it drifts from the published schema is what keeps it honest. Haven't built the fixture yet, so good to have that before I do. And yeah, commentId vs commentIds is exactly the kind of difference a clean rename would quietly wipe out.
"Run them yourself" is the best possible opening — reproducibility is the thing 90% of agent posts are missing, so this already stands out.
The point buried in here that I love: a good mock is documentation, because it encodes the contract the real path is supposed to honor. When the agent path diverges from the mock, the mock isn't wrong — it's the spec catching a regression.
Did you end up promoting the mock into an actual contract test, or keep it as reference? Feels one step away from being a guardrail.
Not yet, it's still just reference. I'd push back a little on the mock being the spec though. A mock only catches regressions if it matches a contract everyone agreed on, and Pouya's had better names than mine that were still wrong for my data. If I'd treated it as the spec I would've locked the mismatch in. A fixture generated from the real schema is where I'd go, and even that only proves the shape matches, not that the authenticated MCP path actually works.
This is the trap with agent integrations, the 120 docs were readable anonymously through the project API so the mock looked right, but the MCP endpoint needing an org token is a different access path entirely. Validating only the public path and merging in 36 hours is how that gap ships. Did the token requirement surface in testing at all, or only when a real agent call hit it?
It showed up in his testing, not mine. Pouya was hitting the 401 on the Context MCP endpoint, so he tested the query generation against an offline mock schema instead, and that's where those field names came from. I hold the token, so on my machine everything just worked, for a reason I never had to think about.
The distinction between reproducing the data and reproducing the actual agent path is really important. I also like the idea of shipping a generated contract fixture instead of leaving contributors to invent their own mock schema.
Yeah, the data being public is what made it look reproducible. I couldn't see the gap myself because I hold the token, so everything just worked for me. A generated fixture is the way I'm leaning, it's not built yet though.
Glad I could help! ✌️