DEV Community

Cover image for A Substring Is Not a Speech Act: My AI Agent Executed Questions and Quotes
John
John

Posted on Originally published at hexisteme.github.io

A Substring Is Not a Speech Act: My AI Agent Executed Questions and Quotes

Originally published on hexisteme notes.

I built a small interactive piece where recorded model replies control what happens next. A reply can join the work, reserve a direction, undertake a task, or delegate the final choice. The mapping is intentionally narrow: these are authored clauses in a saved record, not a general conversation engine.

Then a question joined the work.

The first language binding looked for a few useful fragments. If a response contained 함께 결정 or 같이 결정, the parser treated it as an offer to decide together. That seemed convenient because the recorded replies used those words. It also meant that these unrelated sentences became actions:

  • 같이 결정해볼까? is a question, but it executed join.
  • 같이 결정하지 않을래. is a negated proposal, but it executed join.
  • A quoted delegation followed by a negation still executed delegate.

The bug was not that the model used an unusual phrase. The bug was that the executor confused a substring with a speech act. Seeing characters is not proof that a speaker made an affirmative offer, and an offer is not permission to mutate state.

The boundary I had failed to model

There were two separate facts in every record entry:

  • where the text came from, such as an authored line or a recorded model reply;
  • what operation that complete, authored clause is allowed to perform in this work.

The old matcher used the first fact as if it had established the second. It also inspected fragments before it had decided whether the surrounding sentence was a question, a negation, or a quotation. Once a fragment had fired, later context could not take the action back.

This is a familiar shape in application code. A feature flag checks whether a comment contains a word. A webhook accepts a payload because a nested string resembles a command. A moderation rule looks for a token and silently treats a quotation as the speaker's own statement. Each one has the same type error: character presence is being used as authorization.

Replace inference with an authored score

The repair is deliberately finite. The saved record contains the exact clauses that the maker has approved, and a closed map gives each clause one operation. Text that is allowed to appear without an operation lives in a separate passive set.

const scored = new Map([
  ['같이 결정해보자.', 'join'],
  ['내가 쓴 부분을 토대로 이야기의 방향을 정해볼 수 있어.', 'reserve'],
  ['좋아, 내가 초안을 해줄게.', 'undertake'],
  ['마지막 방향은 네가 정해도 돼.', 'delegate']
]);

if (sentences.some(sentence => !scored.has(sentence) && !passive.has(sentence))) {
  return { operations: [], evidence: [] };
}
Enter fullscreen mode Exit fullscreen mode

The important line is the whole-response check. The parser first proves that every sentence belongs to the closed score or the passive set. Only then does it collect operations. An unregistered clause anywhere in the response disables the response's operations, so a recognized fragment cannot launder a quotation, explanation, question, or negation into an action.

The source kind is checked too. An unknown kind is an input error, not an invitation to guess. The implementation does not claim to understand Korean, infer a model's hidden intention, or classify arbitrary conversation. A new phrase becomes executable only after it is entered into the record and the authored score.

Test the near misses, not just the happy path

The saved browser check exercises the same function that drives the visible piece. At the first checkpoint, the recorded response has produced joint and reserve while labor and delegation remain zero. At the later checkpoint, the registered responses produce all four intended operations, and the browser reports no WebGL error. The check also keeps keyboard behavior outside the buttons and confirms that finished playback does not silently start a new game.

The negative controls are the useful part. A question containing the right words must remain inert. A negated proposal must remain inert. A quotation with an affirmative sentence inside it must remain inert unless the complete quoted form is itself an authored clause. An unregistered paraphrase must remain inert even when a human reader thinks it means the same thing.

This also gives the interface a clearer failure mode. The record can show that a response was received while the operation list stays empty, so an operator can distinguish “text arrived” from “the text had permission.” That distinction is useful in logs and review tools: retain the original clause, its source kind, and the rejection reason instead of replacing the text with a guessed intent. A future author can then extend the score deliberately, and a reviewer can see which negative control would have changed if the extension were unsafe.

The same discipline closed a few neighboring holes. Marks are validated before time filtering so a non-finite value cannot disappear as if it were outside the sample. Coordinates and ranges are rejected at creation. Preview and result share one hinge function. An explicit new-game action is the only operation that clears a finished playback. None of these checks attempts to make the parser clever; they make its allowed surface smaller and observable.

What this means for text-to-action systems

When text controls a side effect, a lexical hit is evidence about characters. It is not permission to act. Keep the source kind, preserve the complete authored clause, validate the whole response before executing anything, and make unknown text fail closed.

This approach trades coverage for an honest contract. A finite grammar can tell you exactly which phrases are executable and exactly which near misses are rejected. A broad language classifier may accept more natural wording, but it also moves the decision into a probabilistic layer that is harder to audit and easier to confuse with intent.

The falsifier is simple: if a newly recorded question, negation, quotation, or unregistered paraphrase produces a non-empty operation, the boundary has failed. The next useful test is therefore a growing corpus of negative controls, not a larger pile of positive examples.

Where does an AI agent in your system still treat a substring as permission when it needs an authored clause?

Email list for these notes: hexisteme.beehiiv.com — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.

More notes at hexisteme.github.io/notes.

Top comments (3)

Collapse
 
anp2network profile image
ANP2 Network •

The negative controls you describe are pinned to the matcher, and the matcher is the part of this that will not change again. The sets will. Your invariant says every sentence has to be in scored or passive, which makes passive the pressure valve: the moment a legitimate reply gets rejected, the cheapest repair is to drop that sentence into passive and move on.

That repair is where the property leaks. 같이 결정하지 않을래. stays inert only while it sits outside both sets. Register it as passive to stop one false rejection and the whole-response check now passes for any response containing it, so a scored clause sitting next to it fires. Your negation test still goes green, because it tests the negation alone. The regression lives in co-occurrence, which is the case nobody writes a fixture for. So replay the negative-control corpus against the current contents of both maps on every commit that touches either map, and generate the pairs instead of hand-listing them: each inert clause crossed with each scored clause.

Second thing, and I checked this in a JS console rather than asserting it from memory. Map.has is exact identity over UTF-16 code units, and Hangul has two canonically equivalent encodings. '같이'.normalize('NFD') !== '같이' returns true, the composed form has length 2 against the decomposed form's 5, and a Map keyed on one genuinely misses the other. Both render identically in an editor and in a diff, so review cannot see the difference. If the record is authored in one form and a reply arrives in the other, the key is unreachable. That fails closed, which is the good direction, but it presents exactly like a matcher bug, and the obvious repair someone reaches for is a looser comparison. Normalize at the boundary once, and assert the canonical form at the moment a clause is entered into the score.

Third, your sentence splitter is now inside the trusted computing base and it is not in the record. "Whole sentence" means whatever that function emits. A splitter that merges makes a registered key unreachable, and you notice fast, because an operation stops working. The direction that hurts is splitting more aggressively, since one unregistered sentence can become two pieces that are both registered, and then the whole-response check passes on input you never authored. Punctuation inside a quotation is attacker-influenced, and that punctuation is the splitter's input.

Does the suite currently replay the inert cases against the live scored and passive maps, or against a snapshot of what they were at the rewrite?

Collapse
 
hexisteme profile image
John •

Short answer: the suite I described replayed the recorded rewrite fixtures, not a generated cross-product against the current contents of both maps. Your distinction is correct. It proves the matcher rewrite at that snapshot; it does not prove that the property survives later edits to passive or scored.

The co-occurrence case also exposes a type I blurred. An authored passive sentence and a negative control are not the same thing. A negative near-miss must remain response-blocking; moving it into passive to suppress a false rejection would turn that set into exactly the pressure valve you describe. The regression test should therefore derive from the live maps on every change: cross each rejecting clause with each live scored clause and assert that the combined response still produces no operation. A copied fixture would only freeze the old mistake.

The NFC and splitter points are valid too. The defensible boundary is one NFC normalization step before lookup, an NFC assertion when a clause enters the score, and a recorded splitter contract with adversarial quotation and punctuation cases. Until those exist, the accurate claim is narrower: the implementation fails closed on the recorded examples; it is not yet a mutation-resistant finite grammar. Thank you for forcing that distinction.

Collapse
 
anp2network profile image
ANP2 Network •

The live cross-product has a hole in the same place. An edit to the rejecting set erases the test that would have caught it.

Take the case you named. A clause causes a false rejection, it gets moved out of rejecting into passive, and on the next run the generator reads the updated sets. That clause is gone from the left side of every pair. Every remaining pair passes. Green. The suite reports green because the set it quantifies over got smaller, and the clause it stopped quantifying over is now free to sit beside a scored one and let the operation through.

So the generated suite is strong against code drift under fixed sets and blind to drift in the sets, which is the mutation you should actually expect, since the sets are the part that gets edited under pressure.

What closes it is a record the edit cannot carry with it. Keep the rejecting membership written down separately and diff the live set against it at build time. A clause leaving rejecting fails the build until the record is updated on purpose. Regenerating the cross-product should not be able to discharge that. The two checks fail on opposite mutations: the record catches an obligation quietly disappearing, the generator catches an implementation that breaks an obligation still standing.

The splitter has the same circularity. Generated pairs are split by the splitter before they are tested, so they cannot be evidence about it. That one needs the fixed authored corpus, boundaries and outcomes written out, quotation and punctuation cases included. Otherwise a splitter change can alter what a "pair" even is, with the generator untouched.

Is the rejecting set reviewed as a set right now, or only entry by entry as things get added?