"Can You Point Us to Clause 23.7?" — The Clause Was Never There
The email landed at 3:04 PM on a Thursday. Subject line: "Question on clause 23.7."
"Could you point us to the exact wording of clause 23.7? We can't find it in the contract we have."
I remember the stomach-drop. I opened the source contract, hit ctrl+F, typed "23.7." Nothing. Typed "23." Still nothing. Scrolled the whole document just to be sure. There was no Article 23, no Section 23.7, no provision even close to the concept the model had attached to that number.
But there it was in our contract risk summary, cited as casually as any real clause. The paragraph before talked about indemnification. The paragraph after talked about limitation of liability. And in between sat a phantom clause number, formatted like it had been there since signing day. I'd skimmed past it during review. So had two other people.
We sent the client's legal team a carefully vague "we're looking into it" and spent the afternoon trying to figure out what the hell had happened. That evening we shipped the first version of source verification. An ugly script, the kind you write when you're embarrassed and in a hurry: pull every clause reference out of the generated text, check it against the source document, flag anything without a match. It printed lines like [verify] clause_ref='23.7' → no match in source_contract_v12.pdf → FLAG and dumped the flagged sentences into a review queue. Over the next few days, it caught a few more phantom references.
The lesson was already clear: a fluent, confident hallucination is dangerous because it looks exactly like the right answer. Internally consistent, well formatted, entirely fictional.
What 23.7 taught us is that the line between a model producing a statement and an organization asserting a fact has to be mechanical. Good intentions don't hold boundaries. Gates do.
Our first gate was format-level. We forced the summary generator to output a strict JSON schema with clause_id, summary_text, and source_ptr. Token masks, grammar constraints, the whole constrained-decoding setup. The output became valid JSON every time. And it did precisely nothing about hallucinated content. A sentence can be grammatically perfect, structurally valid, and completely fabricated. Format is hygiene. Verification belongs somewhere else.
So we moved the gate to the output side. Every assertion in a compliance-sensitive summary now has to carry a source pointer back to a verified document fragment — per clause, per sentence if the text is ambiguous. If a block can't resolve a pointer, it doesn't stream out.
The first implementation used embedding similarity with a cutoff near 0.78. We lost about ten days to that number. It kept rejecting correct summaries that paraphrased the source instead of quoting it, and we kept tuning the cutoff until we realized the real problem: we were using a similarity heuristic to answer a question about identity. So we flipped to a hybrid: deterministic phrase matching against the source fragments first, embedding scoring as a second pass. Over the next several months, that layer caught eight hallucinations that would have changed a downstream decision. The people who read those summaries never saw them. The audit log shows the rejections, and a person confirmed each one as a genuine error.
Then came the consistency problem. Run the same request twice with different random seeds, measure how much of the output overlaps. We call that metric "reproducibility," and we deliberately keep it out of any UI where someone might read it as confidence. High overlap only means the model reliably tells one stable story. It can be stable and false. Low overlap is the operational signal: it marks the case as ambiguous enough to route to a person. It doesn't certify anything.
The full audit trail was the most expensive part to build. After 23.7, we wanted to replay every generation: which contract version, which retrieval fragments, which decoding parameters, what the verifier said, what the reproducibility score was. Every release point that can affect an obligation — a permission grant, a coverage decision, a trigger condition — records those fields. If a decision can't be rebuilt from the log, the audit trail is decoration.
The costs are real. Constrained output reads stiffer. That's the intent. Only high-risk acceptance gates run the full cage; ordinary chat keeps its fluency. Verification adds latency: a fully source-verified summary takes nine seconds where an unverified one took two. We run the full path on anything headed to a client and route casual internal queries through a lighter heuristic filter. Multiple sampling runs multiply inference cost. We treat it as insurance for the subset of decisions that can trigger real-world side effects. The rest of the traffic doesn't pay for it.
One operational detail nobody warns you about: the similarity cutoff started out in a runtime config file. A teammate adjusted it one afternoon while debugging something unrelated, and for the next ten days, summaries went out with thin sourcing. Nobody noticed until we pulled the audit. The cutoff now lives in the deployment artifact. Changing it requires a release. If the organizational layer can't hold a boundary, the algorithmic layer won't either.
If this stack had existed when the 23.7 summary was generated, that citation would never have reached the email. The source pointer would have found nothing. The verifier would have rejected the sentence. The text would have been rephrased or sent back to the editor. The model would not have been asked to explain itself. It would have been treated as suspect until it pointed to a source. Trust lives in the workflow, in the mechanical parts around the model.
These days, every conclusion an AI produces in a compliance pipeline has to point back to a source. If it can't, it should not be believed. That's the whole job.
Top comments (0)