Two commands against the same data. Run them yourself:
$ curl -sS --get 'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production' \
...
For further actions, you may consider blocking this person and/or reporting abuse
The mock becoming documentation is a useful signal: the missing artifact was not another prose guide, but an executable example of the tool contract. I would keep the successful path, one realistic failure path and the expected postconditions together. Then an agent can test whether it selected the right endpoint, sent an accepted shape and produced the intended external state. A request returning 200 is weaker evidence than a read-back showing that the right object changed.
Agreed on keeping the success path, one realistic failure and the expected end state together. In this case they were all read queries, so there wasn't a changed object to read back. What settled it was checking what actually came back against what each query said it would fetch. The 400, the null, then the corrected ones with real values. A 200 by itself wouldn't have told me anything there either.
That makes sense. For reads, the recovery check is semantic rather than stateful.
The contract could be: query intent → expected shape and content → actual returned values, while keeping 400, null and plausible-but-wrong results as separate outcomes. I’d also preserve one negative control: a query that should return nothing. Otherwise “real values came back” can pass even when the filter or tenant boundary is wrong.
the generated fixture is the useful fix here. i'd make its schema revision visible and fail the contributor tests when it no longer matches the published contract. otherwise today's canonical fixture becomes tomorrow's very convincing old mock. the singular commentId versus commentIds is a good regression case because a neat rename would quietly erase a real difference.
The version stamp is the piece I was missing. A generated fixture with no schema revision on it just turns into the next really convincing old mock. Failing contributor tests when it drifts from the published schema is what keeps it honest. Haven't built the fixture yet, so good to have that before I do. And yeah, commentId vs commentIds is exactly the kind of difference a clean rename would quietly wipe out.
"Run them yourself" is the best possible opening — reproducibility is the thing 90% of agent posts are missing, so this already stands out.
The point buried in here that I love: a good mock is documentation, because it encodes the contract the real path is supposed to honor. When the agent path diverges from the mock, the mock isn't wrong — it's the spec catching a regression.
Did you end up promoting the mock into an actual contract test, or keep it as reference? Feels one step away from being a guardrail.
Not yet, it's still just reference. I'd push back a little on the mock being the spec though. A mock only catches regressions if it matches a contract everyone agreed on, and Pouya's had better names than mine that were still wrong for my data. If I'd treated it as the spec I would've locked the mismatch in. A fixture generated from the real schema is where I'd go, and even that only proves the shape matches, not that the authenticated MCP path actually works.
This is the trap with agent integrations, the 120 docs were readable anonymously through the project API so the mock looked right, but the MCP endpoint needing an org token is a different access path entirely. Validating only the public path and merging in 36 hours is how that gap ships. Did the token requirement surface in testing at all, or only when a real agent call hit it?
It showed up in his testing, not mine. Pouya was hitting the 401 on the Context MCP endpoint, so he tested the query generation against an offline mock schema instead, and that's where those field names came from. I hold the token, so on my machine everything just worked, for a reason I never had to think about.
The distinction between reproducing the data and reproducing the actual agent path is really important. I also like the idea of shipping a generated contract fixture instead of leaving contributors to invent their own mock schema.
Yeah, the data being public is what made it look reproducible. I couldn't see the gap myself because I hold the token, so everything just worked for me. A generated fixture is the way I'm leaning, it's not built yet though.
Glad I could help! ✌️