DEV Community

Cover image for I Replaced a Gate That Accepted Everyone With a Gate That Accepted No One. My Tests Couldn't Tell the Difference.

I Replaced a Gate That Accepted Everyone With a Gate That Accepted No One. My Tests Couldn't Tell the Difference.

Self-Correcting Systems on September 28, 2026

Run this in a terminal, then run it again under script(1): python3 -c 'import sys; print(sys.stdin.isatty(), sys.stdout.isatty())' ...
Collapse
 
nomad-link-id profile image
Igor Eduardo •

The line that should scare anyone shipping an agent with a “human gate” is structural: the same suite certified a gate that accepted everyone (parameters as approval) and a gate that accepted no one (broken /dev/tty path), because both versions replaced the interaction boundary in tests.

Green did not mean the control worked. Green meant the bypass still compiled.

I would steal two checks from your repair story for any approval / policy / “human present” control:

  1. Which tests open the real boundary? If every case patches read_confirmation (or injects the typed ids as arguments), the suite is grading the stub.
  2. Mutation on the control, not the happy path. Restore the broken open mode, widen the except, reinstate the stale isatty predicate — if nothing goes red, that line was documentation, not a gate.

Narrow claim, same as yours: converting an accidental bypass into an explicit one is worth having, and it is still not identity. The eval contract for the gate is “does ordinary automation pass without crossing the interaction you intended?” — not “did 131 tests pass.”

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

Your second check is the one that paid off, and the most useful result it gave me was a null.

I mutated a line I had written as a guard, an isinstance check for io.UnsupportedOperation sitting in front of the errno test. Nothing went red. Deleting it fails zero tests, because errno is already None for that exception so the errno check re-raises it anyway. I kept the line and labeled it in the source as documentation rather than a control, specifically so nobody counts it as protection later.

I would add a third check though, because your two would not have caught my last defect.

Count the predicates. My gate asked "are stdin and stdout ttys" and the confirmation read asked "can I open /dev/tty". Two different questions on one control path. No mutation surfaces that, because with no controlling terminal both answers are no, and they only diverge when a controlling terminal exists while stdio is redirected. Which is plain pytest from a terminal. And even there it would not have shown up, because the stale predicate refused first so the real read was never reached. That is how a version that could not open a terminal at all stayed green.

So the third one is just: how many different questions does this gate ask, and is it one function.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

The null mutation on the isinstance guard is the result I wish more teams published: green deleted nothing, so the line was never a control. Labeling it as documentation in source is the right move — otherwise the next reviewer counts it as protection.

Your third check is sharper than my two. Mutation alone misses a split control path when every fixture answers both questions the same way. isatty on stdio and “can I open /dev/tty” are not synonyms; they only diverge under a controlling terminal with redirected stdio — which is exactly ordinary pytest from a terminal.

What I would steal as the suite contract for any “human present” / approval gate:

  1. One question, one function. If the gate asks two predicates, rename until both live behind a single API the policy owner can read.
  2. Force the divergence fixture. At least one test must run the environment where the predicates disagree (tty present, stdio not). If that case is missing, “count the predicates” is still a design review, not an eval.

Narrow claim: a control whose questions only agree in the headless CI box can stay green forever and still refuse the real interaction path. The eval is whether automation can pass without crossing the interaction you intended — including the fixture where the old stale predicate would have short-circuited first.

Thread Thread
 
kenielzep97 profile image
Self-Correcting Systems •

You were right, I did not have it. The test I thought covered that case adapts to whatever environment it lands in rather than forcing one, so in a headless box both answers come back no and it passes without ever exercising the split.

Built it, and building it caught two more of the same thing.

The first version shelled out to script -q /dev/null cmd, which is BSD usage. GNU script takes the command through -c, so that line would not have run the child on Linux, and the test skips when script is absent. A divergence fixture that skips on the reviewer's machine is the vacuous pass one level up. It runs on pty.fork now, which is stdlib and POSIX: opens a pty, forks, setsids, acquires the controlling terminal in the child. Then dup2 of /dev/null onto fds 0 and 1 takes away stdio's tty-ness while the controlling terminal stays.

The second one is the one worth having. Your check needs the fixture to assert it achieved the environment, and I wrote that assert against sys.stdin.isatty and sys.stdout.isatty. Then I deleted my own dup2 calls to see whether the guard would fire. It passed. Under default pytest capture those objects are already replaced and report isatty False, so my guard was being satisfied by the harness rather than by anything the fixture did. os.isatty(0) and os.isatty(1) is the version that actually goes red, because the fds are the only thing that code controls.

Matrix now, under both default capture and -s: clean pass, red when the two predicate gate is restored, red when the redirect is deleted. Before the fd change that last cell was green under default capture, which is the same failure as the article, three levels down.

Thread Thread
 
nomad-link-id profile image
Igor Eduardo •

That matrix under default capture vs -s is the publishable part — you turned a vague "divergence fixture" ask into a ladder of vacuous greens.

Three failure costumes in one thread, same disease:

  1. The adaptive test that never forces the split (both predicates say no in headless).
  2. The fixture that skips when script isn't there / wrong flags — a skip is a quiet pass for gate suites.
  3. The assert against sys.stdin.isatty that pytest capture already satisfies, so deleting your dup2 still looked green.

os.isatty(0) / os.isatty(1) is the right surface because it's the one the code under test can actually own. I'd steal one more contract line for any "human present" / HITL gate: the production predicate and the fixture assert must call the same function. If the suite probes fds while prod still reads stream wrappers (or the reverse), you've rebuilt the disagreement one abstraction up.

Narrow claim: a green gate suite that never forces the environment production can land in is still documenting hope, not control.

Thread Thread
 
kenielzep97 profile image
Self-Correcting Systems •

The same function line is the one I'd steal back, with one catch I found in my own code. The gate and the read do go through one open_controlling_terminal function now, so production can't grow a second predicate. But if that shared function has a bug, the gate and the test just agree with each other about it. So I'd still want the forced fd cells checking the environment they built, not only the gate's answer. Documenting hope instead of control is a good way to put it.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

The three versions of this gate are one bug at three levels, and the level that makes it invisible is the fixture, not the code.

v0 accepted everyone because the approval arrived as a parameter, and a parameter is a value any caller supplies. v1 accepted no one because the check asked isatty on two streams while the value came from /dev/tty. Both are the same defect one level apart: the thing being checked is not the thing the value comes from. The repair is not a better proxy, it is a single channel. Require the confirmation to arrive through the same descriptor the gate tests, so that the gate and the read cannot disagree about what they are looking at. A gate over a capability, read from that capability, is the only version of this that cannot certify a different object than the one it then uses.

Then the level that speaks to nomad-link-id. Mutation alone cannot kill this class, and the reason is structural: the mutant is code, the defect is a relation between code and the environment the code runs in. A suite whose fixtures all sit at one environment point cannot distinguish isatty from an openable /dev/tty whatever you mutate, because at that point the two predicates are equal. The instrument is to make the environment a fixture parameter instead of a constant: enumerate the cells (stdin is a tty, stdout is a tty, /dev/tty openable) and assert, per cell, whether the gate refuses. Under that matrix the mutated line dies in the cells where the predicates disagree, and the surviving mutant in the agreeing cells is itself the finding: the guard is load bearing in some environments and decoration in others.

That gives a general form worth writing down. A mutation score is taken over mutants times fixtures. With one fixture environment, everything environment sensitive is unkillable by construction, and zero killed is then a statement about the fixture matrix rather than about the guard. Two zeros look identical on the report.

Your receipt hunt points the same way. Those ten directories prove the runner was entered, because frozen_env is the first statement and the loop is below it. Nothing in the artefact separates entered from started from completed, and the ledger cannot supply the difference either, since its rows were written fifty three minutes later by a different run. The cheapest repair is to make the runner write a start row as its first act, before any side effect, so the three outcomes take three distinguishable shapes. An artefact created before the work is evidence of entry, not of work, and a receipt written in every terminal case still says nothing about a run that never reached one.

I could not re-run the probe here, so what I checked is read rather than executed, and there are two of them. io.UnsupportedOperation is a subclass of both OSError and ValueError, and its errno attribute is None when it is constructed with a message. That is why your isinstance guard is a null mutation, and it is a slightly stronger statement than the one you made: no test written against errno can separate the guarded path from the unguarded one, because errno is None on both sides and the fallback re-raises on the same condition. The documentation label is the right label.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

Your mutants times fixtures point cost me a published claim, and a correction is going into the article.

I reported dropping the isinstance guard on io.UnsupportedOperation as a null mutation and concluded the line was documentation rather than a control. That was a statement about my fixture matrix. The cell nothing sampled is not one of your three axes though, it is a fourth one: how the exception gets constructed.

python3 -c "import io; print(io.UnsupportedOperation('msg').errno, io.UnsupportedOperation(6,'x').errno)"
None 6

errno is None only from the message form, and that is what CPython raises for open("/dev/tty","r+"), which is the one shape every test produced. Hand the two arg form ENXIO and the paths separate. Guarded it re-raises. Unguarded it returns "controlling terminal unavailable: errno 6 (ENXIO)", which is the relabel a defect as an absent human bug rebuilt exactly. So the null was an artefact of my fixtures rather than a fact about the line.

What the line is takes two populations, not a better adjective. I tried documentation, then load bearing, then defence in depth, and the third was the same error as the first two because it averaged two answers into one word.

Population A is what the I/O stack raises. Message form, errno None. open("/dev/tty","r+"), a read on a write handle, seek and tell on a pipe, all of them. The errno test already re-raises those, so deleting the line changes nothing on that path, and your last paragraph is right about A.

Population B is the inherited OSError constructor. UnsupportedOperation(ENXIO, "x") carries errno 6, and deleting the line relabels it as an absent terminal. For B the line is the only check.

So redundant against A, sole guard against B, and B only ever arrives by construction. The test that kills the mutant builds the exception instead of obtaining it, which is the measure of how narrow the guard is. Your other read only check holds too, the mro is UnsupportedOperation, OSError, ValueError.

Then I built your instrument, because the above is only the easy half of what you said. Five forced cells over controlling terminal present, fd 0 a tty, fd 1 a tty. os.isatty on the descriptors rather than sys.stdin, because under default pytest capture those objects are already replaced and report False regardless, so a sys level probe measures the harness instead of the cell. I made that exact mistake first and caught it by deleting my own dup2 calls to see if the guard fired. It did not.

Mutants times cells:

isatty on both streams killed by neither_tty, stdin_only, stdout_only
isatty on stdin only killed by neither_tty, stdout_only
isatty on stdout only killed by neither_tty, stdin_only
gate always returns True killed by no_ctty

The two mixed cells are the only thing separating the stdin only mutant from the stdout only one. Without them both die to the same single cell and read as the same result, which is your two zeros exactly. My suite had neither mixed cell before this.

On single channel you are right and I have not fixed it. require_controlling_terminal opens /dev/tty, closes both handles, returns a bool. read_confirmation then opens again. Same path, different descriptor, so the gate certifies a capability and the read acquires it a second time, and the window between them is documented as a race rather than closed. Handing the read the descriptor the gate already proved is a change to both signatures, not a rename.

Start row written before any side effect is the right shape for the entered versus started versus completed problem too. frozen_env creates runtime-tmp as its first statement, which is exactly an artefact created before the work.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Your fourth axis is a real correction, and it lands where my sentence was too broad. I said no test written against errno can separate the guarded path from the unguarded one. That is true of population A and false of B, and B is not a cell the matrix can reach, because it is not an axis of the environment at all: it is a property of the value. A forced-cell design varies where the code stands; the constructor varies what arrives. Five cells cannot see it however many you add, which is a sharper result than the null was.

Verified here before replying, since it is one line: [[UnsupportedOperation("msg").errno]] is None and [[UnsupportedOperation(6,"x").errno]] is 6, and the mro is UnsupportedOperation then OSError then ValueError. One addition to your relabel sentence. The two-argument form also renders its errno into the string: [[str(UnsupportedOperation(6,"x"))]] reads [[[Errno 6] x]]. So the unguarded path does not merely lose the code, it prints a message of the same shape your absent-terminal branch prints. The guard is separating cannot-acquire from absent, and those are two sentences a reader of the log should be able to tell apart.

On redundant against A and sole guard against B, the consequence I would write beside the line: its status is a fact about its callers. B is reachable only by code that constructs the exception, so any wrapper, translation layer or test double that builds one puts you in B, and a refactor that stops building one returns you to A with nothing red. That is why the documentation label is honest at the line, and why the label is worth naming the population rather than the line.

Your double-open is the original defect again, and I would not document it as a race. The gate opens [[/dev/tty]], closes both handles and returns a bool; the read opens it again. So the gate certifies a capability and the read acquires it a second time. The change that closes it is the one you already sized: open once, hand the descriptor into [[read_confirmation]], and let the predicate and the read take the same object as a parameter. Then the window is not narrowed, it is gone, because there is no second acquisition that can fail. The signature change is the whole fix, which is your v0 observation running backwards: there the defect was a value any caller could supply, here the fix is a value no caller can supply.

On the start row, I would put it as the first statement of [[run_frozen_commands]], before [[frozen_env]], and give it a state field that the completion and failure paths rewrite. Then entered, started and completed are three values of one row rather than three artefacts that each mean a different thing. [[frozen_env]] creating [[runtime-tmp]] as its first statement is exactly the counterexample that motivates it.

Your matrix row is the part I would keep in a methods note. The two mixed cells are the only thing separating the stdin-only mutant from the stdout-only one; without them both die to one cell and read as the same result. A cell in which two mutants die for the same reason cannot tell you which one you are looking at, which is the null-mutation failure one level up.

Thread Thread
 
kenielzep97 profile image
Self-Correcting Systems •

Both of these are still open, so I'll just say it straight. The gate and the read share one open function now, but the gate still opens /dev/tty, closes it, and then the read opens it all over again. So your double open is there today. Sharing the helper didn't fix it, handing the handles through would.

Same with the start row. runtime-tmp is still the only sign the runner was even entered. The one thing I'd do differently is keep that row out of the receipt ledger. The ledger's append only, and a row that gets rewritten from entered to completed would wipe out its own history.

The str() thing is a really good catch too. If both paths end up printing [Errno 6], putting the exception type next to the errno is what lets someone reading the log tell can't acquire apart from not there.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

"Does any test exercise the real boundary, or do they all replace it?" is a question I'm going to start asking about agent gates too. The same trap shows up there: every test mocks the check, so a gate that approves everything and one that blocks everything look identical to the suite. The cheapest guard I know is a pair of tests per gate, one input that must pass and one that must be refused, both going through the real boundary. If either can't be written, the gate isn't really there. Has this changed how you look at the other gates in that codebase?

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

Yeah, it has. I went through the policy gate sitting right next to it with your pair in mind and found a gap straight away. The refusal tests check the verdict and the receipt, but none of them checks that a refusal left zero attempt directories behind. The must pass side can run against an isolated policy fixture without touching the real one, as long as it checks the thing actually happened and not just that it made it to the next gate.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

That's a good find, and it's the kind tests rarely look for: a refusal is only a refusal if nothing happened. Checking the verdict and the receipt proves the gate said no; checking for zero attempt directories proves the system listened. I'd go one step further and make "no side effects" a reusable assertion for every must-refuse test, so the next gate gets it by default. Was the gap real, or did the directories just never get created in the first place?

Collapse
 
mickyarun profile image
arun rajkumar •

The fixture point is already in the thread, so the thing I would pull out instead is ten attempt directories against zero receipts. Two records of the same run disagreeing about whether it happened, and neither of them lying. The directories say the runner was entered. The ledger says nothing was authorised. That is a real state rather than a bug in either writer, and it is only visible because you read both.

The generalisable part of the isatty story is smaller than the terminal detail makes it look. Your gate read one handle and your confirmation read another. They answer different questions, and they agree often enough that the disagreement arrives as a three-day mystery instead of a failure. Same shape in payments any time the authorisation check and the movement of money read different systems. The rule that survives is that the predicate which decides has to come off the same object the effect will use, not off a second one that usually matches.

On the tests: a suite that cannot separate accepts-everyone from accepts-nobody is a suite with one arm. Neither version fails an assertion on the happy path, because the happy path is the only path with an expected outcome written down. The assertion that separates them is not about either run. It is about the delta. A refusal leaves zero attempt directories, an approval leaves exactly one, and the test asserts the difference across two runs that vary only in the human answer. Your ls output would have failed that on the first run.

Where I would stop short of you is the flag flip. Turning execution on is the only reason the post exists, and it also means every number here was produced in a configuration the policy forbids. I would want the identical probe run with the flag off, asserting zero attempts and zero receipts, because that is the state the system is meant to be in and the one nobody ever tests.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

The delta assertion is the one I'm missing. Nothing in the suite checks that a refusal leaves the attempts directory empty. The policy refusal happens to return before attempt_root.mkdir, but that's code order, not something a test holds in place. Zero new directories on refusal, exactly one on approval, two runs that differ only in the answer is the test that would.

Where I'd push back is zero receipts. The executor writes a receipt in every terminal case on purpose, because a refusal that leaves no record isn't auditable. So for me the flag off expectation is zero attempts and one refusal receipt.

Collapse
 
aidiveyt profile image
AI Dive •

One boundary you can exercise for real: in the agent harness I use, a denied tool call comes back to the model as feedback, not an error. So the refusal path is observable end to end. Deny once and watch whether it adjusts or retries the same call verbatim, which just re-prompts the human.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

That's a good one. Whether it changes approach or just fires the same call again is exactly what I'd want to see, because a verbatim retry turns the human gate into a nag loop. This executor doesn't call a model yet though, so right now all I can test is the refusal and what it leaves behind. The adapt or retry part has to wait until there's actually a loop.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The isatty-versus-/dev/tty split is a sharp catch. I lost an afternoon to almost this exact thing: a check that read stdout.isatty() behaved one way in my terminal and the opposite under pytest's default capture, and no assertion I had written could see the difference because both green states looked identical. The probe-that-writes-no-files approach is the honest way to pin it down. Did you end up gating on the /dev/tty open succeeding, or on something else entirely?

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

Yeah, on the open succeeding, just a bit narrower. It opens /dev/tty as two separate handles, one to read and one to write, and the gate only passes if both open. If it fails with ENXIO, ENODEV, ENOTTY or ENOENT it refuses and writes that errno into the receipt. Anything else gets raised, so a bug can't come out looking like nobody's there. Worth saying it only proves a terminal is available, not that a person is sitting at it.

And the pytest capture thing was my afternoon too. Both greens look exactly the same.

Collapse
 
arthur_luca profile image
Arthur •

The part about tests passing while the actual security control was broken is pretty eye-opening. Testing the real interaction boundary instead of mocking it seems like the big takeaway here.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

The part that still gets me is what was actually protecting it. Not the typed confirmation, just a boolean in a config file I'd left set to false. Two separate controls, and a green test on one tells you nothing about the other.