Run this in a terminal, then run it again under script(1):
python3 -c 'import sys; print(sys.stdin.isatty(), sys.stdout.isatty())'
If your code decides whether a human is present by calling isatty(), and then reads that
human's confirmation from /dev/tty, those are two different questions, and there are
processes that answer them differently. This is the story of finding that out three times in
one file, each time because the previous fix was wrong in a way my tests could not see.
Here is the probe. Standard library only, writes no files and mutates no state; full source at the end.
$ script -q /dev/null python3 gate_probe.py --redirected
isatty(stdin) and isatty(stdout) : False
/dev/tty openable : True
detail : opened
open("/dev/tty", "r+") -> UnsupportedOperation: File or stream is not seekable. (errno=None)
THE TWO PREDICATES DISAGREE for this process.
A gate on isatty and a confirmation read on /dev/tty will not agree here.
That process has an openable controlling terminal even though its redirected stdin and stdout
are not TTYs. For three days that was the shape of my gate.
One precision, because a reader will check. Default pytest capture replaces both streams,
which I had wrong in an earlier version of this paragraph. On pytest 9.0.2, launched from a real
terminal:
default capture stdin=DontReadFromInput (isatty False) stdout=EncodedFile (isatty False)
pytest -s stdin=TextIOWrapper (isatty True) stdout=TextIOWrapper (isatty True)
/dev/tty under default capture: OPENS
So plain pytest from a terminal is the disagreement, on its own — no wrapper needed. My gate
required stdin.isatty() and stdout.isatty(), false under default capture for two reasons at
once, while the confirmation path could open the terminal the whole time. Under pytest -s the
two predicates agree, so don't look for a mismatch there.
The system, briefly
An agent that can run three frozen commands against an exported copy of a repository. It has
never been authorized for normal execution — the policy file forbids it — and I temporarily
flipped that flag during the test run described below, which is how the rest of this post
exists.
$ python3 -c "import sys; sys.path.insert(0,'.')
import stage2a_preflight as PF; print(PF.check_policy())"
(False, 'policy forbids execute_shell; owner has not authorized Stage 2a')
Before it runs anything, a human is supposed to be shown what will happen and type two
identifiers back.
v0 — the gate was a parameter
def execute(job, approval, typed_job_id=None, typed_procedure_id=None):
ok, note = confirm_approval(approval, typed_job_id, typed_procedure_id)
if not ok:
return "REFUSED_APPROVAL_MISSING", note
Read it and it looks like a human gate. typed_job_id is a parameter, and a parameter is
something any caller supplies. A test fixture is a caller with two correct strings.
I flipped the policy flag that permits execution and ran the suite. Afterward:
$ ls -lT state/attempts | awk 'NR>1 {print $6, $7, $8, $10}'
Sep 24 16:36:41 TEST-JOB-1-1790282200
Sep 24 16:36:42 TEST-JOB-1-1790282202
Sep 24 16:36:44 TEST-JOB-1-1790282204
Sep 24 16:36:45 TEST-JOB-1-1790282205
Sep 24 16:37:42 TEST-JOB-1-1790282262
Sep 24 16:37:43 TEST-JOB-1-1790282263
Sep 24 16:37:45 TEST-JOB-1-1790282265
Sep 24 16:37:46 TEST-JOB-1-1790282266
Sep 24 16:36:43 TEST-JOB-2-1790282203
Sep 24 16:37:44 TEST-JOB-2-1790282264
Sorted by name, not time. Read the clock column and it is two runs of five, about a minute
apart — 16:36:41–45 and 16:37:42–46 — eight for one fixture job and two for a second. Each
directory holds a 266,240-byte tar of my own repository, all ten identical in size, extracted
into a 27-file tree.
The gate opened exactly as written.
What those ten do and do not prove
They contain a third entry, runtime-tmp, and it is empty. Its existence is not proof the
three commands started, because of where it is created:
def frozen_env(attempt_root):
"""addendum v1 B1. An allowlist, not a filtered copy. HOME absent. No PYTHON*."""
tmp = attempt_root / "runtime-tmp"
tmp.mkdir(parents=True, exist_ok=True)
def run_frozen_commands(export_root, attempt_root):
env = frozen_env(attempt_root)
results = []
for entry in PF.FROZEN_COMMANDS:
frozen_env() is the first statement of run_frozen_commands(), and the for loop below it is
where commands actually launch. So the directory is created before anything runs. So runtime-tmp proves the runner was entered. Whether any command
started is not established by these directories, and I am not going to round that up.
What the ledger says is stranger. The ten exports left no receipt at all:
$ python3 -c "
import json
for l in open('state/receipts.jsonl'):
r = json.loads(l)
print(r['final_state'], r['job_id'], len(r['commands'] or []), r['written_at'])"
REFUSED_NO_HUMAN_PRESENT tg_912616161 0 2026-09-24T21:29:28.666303+00:00
REFUSED_NO_HUMAN_PRESENT tg_912616161 0 2026-09-24T21:30:40.369098+00:00
REFUSED_NO_HUMAN_PRESENT TEST-JOB-1 0 2026-09-28T01:31:55.420602+00:00
Three rows, zero commands. None of them is the export run. Rows 1 and 2 land fifty-three
minutes after it; row 3 is from this repair session, below.
There is also a quarantine file from that day, receipts_TEST_POLLUTION_QUARANTINED_2026-09-24,
and the export run is not in that either — its ten rows are all REFUSED_COMMAND_BOUNDARY
written between 15:02 and 16:06 UTC, while the exports are 20:36-20:37 UTC. A different event,
four to five hours earlier.
So: ten real exports in production state, and no receipt for them in either ledger. Where that
receipt went I cannot establish. execute() writes one in every terminal case, so either it was
written to a redirected path and discarded with a temp directory, or something raised before the
write. I did not preserve the test configuration from that run, and although the directory is a git
repo, those files are not tracked in it, so absence here is absence — not evidence of a
destination.
Those rows read REFUSED_NO_HUMAN_PRESENT, which is the same overclaim I rename a function for
further down. That value is now REFUSED_CONTROLLING_TERMINAL_UNAVAILABLE and the receipt schema
went v1 -> v2 to say so. The three rows above were not rewritten — an append-only ledger
edited to match new vocabulary is not an audit trail — so the file holds both values, and
schema is what tells a reader which vocabulary a row was written under.
The thing that was actually protecting me
Not the typed confirmation. A boolean in a config file that I had left set to false.
v1 — the fix could not open a terminal
Move the read off the parameter list and onto the controlling terminal. There is no argument
to fill:
def read_confirmation(approval):
with open("/dev/tty", "r+") as tty: # <- this line
...
On this macOS machine, open("/dev/tty", "r+") fails on both CPython 3.9.6 and 3.13.9. The
update-mode I/O stack buffers through BufferedRandom, which wants a seekable raw stream; this
terminal is not one:
$ script -q /dev/null python3 gate_probe.py </dev/null
isatty(stdin) and isatty(stdout) : True
/dev/tty openable : True
detail : opened
open("/dev/tty", "r+") -> UnsupportedOperation: File or stream is not seekable. (errno=None)
The last line is the probe opening the same terminal it just opened successfully, in the mode
I had used. Reproduced identically on CPython 3.9.6 and 3.13.9. And then the detail that turned a broken
line into a broken control:
$ python3 -c "import io; print(io.UnsupportedOperation.__mro__)"
(<class 'io.UnsupportedOperation'>, <class 'OSError'>, <class 'ValueError'>, ...)
io.UnsupportedOperation subclasses OSError, and carries errno=None. My handler was:
except OSError as exc:
return "REFUSED_NO_HUMAN_PRESENT", "no controlling terminal"
So a human sitting at a real terminal was told they were not there, and the receipt recorded
the wrong cause. It failed closed, which is the good direction to fail — but the gate now
refused the only caller it was built for.
In v1, every test that exercised approval replaced read_confirmation by name. Not one
exercised the real tty read. v0 had the earlier version of the same blind spot: there was no
read_confirmation to replace, because its tests supplied the ids directly as arguments.
Different bypass, same hole — neither broken version had a human-interaction boundary under test,
so the suite was green for both. It was never testing the gate; it was testing the bypass.
v2 — and the third defect, which the third test found
Separate handles, and an except narrow enough to mean something:
CTTY_UNAVAILABLE_ERRNOS = frozenset({errno.ENXIO, errno.ENODEV, errno.ENOTTY, errno.ENOENT})
def open_controlling_terminal():
out = open("/dev/tty", "w")
try:
inp = open("/dev/tty", "r")
except BaseException:
out.close()
raise
return out, inp
Then a test I wrote for this failed, and I nearly patched the test. That would have been the
fourth version of the same mistake. What it had found:
require_human() asked "is sys.stdin a tty?"
read_confirmation() asked "can I read /dev/tty?"
Two different predicates for one control, and the stale one ran first. They are not
ordered — one asks whether two particular streams are terminal devices, the other asks for the
process's controlling terminal. In the captured pytest configuration I reproduced, isatty
was false while /dev/tty stayed openable, so the stale check refused before the confirmation
path could run. Fail-closed, but the real control was shadowed by a different predicate.
I first wrote that this shadowing was why the r+ version survived the suite. That is a
wrong-reason claim in a post about wrong-reason claims, and the probe disproves it:
no controlling terminal -> OSError errno=6 (ENXIO)
terminal attached -> UnsupportedOperation (errno=None)
r+ fails two different ways, and except OSError swallows both. Delete the shadow and
run again: with no terminal you get a real ENXIO, which is the verdict every refusal test
expects; at a terminal you get UnsupportedOperation relabelled into the same refusal. Either
way, green.
So the shadowing is a second independent reason nothing could have caught it, not the cause.
The causes are the two already named: every test patched read_confirmation, and the except
was wide enough to swallow a defect. What the shadow did do is stop the real path from being
exercised by that integration route — which is why the fix needed a test that calls it
directly.
So there is now one predicate, in one function, used by both:
def require_controlling_terminal():
try:
out, inp = open_controlling_terminal()
except OSError as exc:
return False, controlling_terminal_unavailable(exc)
out.close()
inp.close()
return True, "controlling terminal present and openable at /dev/tty"
And controlling_terminal_unavailable() re-raises anything outside the set, so a programming
fault can no longer be reported as an absent human. The refusal carries the evidence rather
than the conclusion:
controlling terminal unavailable: /dev/tty open failed with errno 6 (ENXIO): ...
Two renames went with this, both the same correction. require_human() asserted something no
check in that file can establish — it cannot prove a person is present, only that confirmation
is obtainable from a terminal, and naming it after the stronger claim is what licensed a second
predicate to grow beside it. And the errno bucket was called NO_CTTY_ERRNOS while containing
ENOENT (/dev/tty does not exist here) and ENOTTY (that fd is not a terminal). Neither
literally means "this process has no controlling terminal." A set named for its strongest member
is the same overclaim one level down, so it is CTTY_UNAVAILABLE_ERRNOS now. The bucket is a
decision about what to refuse on; the errno is the fact.
Passing tests prove nothing, so I broke it on purpose
Four mutations against the gate, each run as the full 131-test suite:
| mutation | caught by |
|---|---|
restore open("/dev/tty", "r+")
|
4 tests |
treat any OSError as an absent human |
3 tests |
restore the isatty predicate |
9 tests and subtests |
drop the explicit isinstance(io.UnsupportedOperation) guard |
nothing. 131 passed |
The fourth row is the useful one, and I drew the wrong conclusion from it. Correction, added
2026-09-28 after this post was published.
I originally wrote that the line was redundant, since errno is already None for
UnsupportedOperation so the errno test re-raises it anyway, and I labelled it in the source as
documentation rather than a control. A reader named howcani
pointed out why that inference does not hold:
A mutation score is taken over mutants times fixtures. With one fixture environment,
everything environment sensitive is unkillable by construction, and zero killed is then a
statement about the fixture matrix rather than about the guard. Two zeros look identical on
the report.
They were right, and the cell my fixtures never sampled is one line of Python:
$ python3 -c "import io; print(io.UnsupportedOperation('msg').errno, io.UnsupportedOperation(6,'x').errno)"
None 6
UnsupportedOperation carries errno=None only when built from a message, and that is the form
CPython raises for open("/dev/tty","r+") — the single shape every test produced. Built with two
arguments it carries an errno. Hand it ENXIO and the two paths separate:
guarded -> re-raised, a defect surfaces as a defect
unguarded -> "controlling terminal unavailable: errno 6 (ENXIO)"
That second line is the relabel-a-defect-as-an-absent-human bug rebuilt exactly. So the null
result described my fixtures rather than the code, and "a line the suite cannot distinguish from
its own absence" was a claim about the suite that I stated as a claim about the line.
And then a second correction on top of the first, because the line has two answers and needs
both stated. I tried "documentation", then "load-bearing", then "defence in depth, constructible
but unreached" — and the third was the same mistake as the first two, since it averaged two
populations into one adjective. An independent reviewer caught that, which is the identical defect
as two zeros looking identical on a report.
Population A — what the I/O stack actually raises. The message form, errno=None. Verified
against open("/dev/tty","r+"), a read on a write handle, and seek and tell on a pipe: all
message form, all None. The errno test below already re-raises those, so deleting the line
changes nothing on that path. For A the line is documentation, exactly as first written.
Population B — the inherited OSError constructor. io.UnsupportedOperation(ENXIO, "x")
carries errno=6, and deleting the line relabels it as an absent terminal — the original bug
rebuilt. For B the line is the only check.
So: redundant against A, sole guard against B, and B arrives only by construction — the test that
kills the mutant builds the exception rather than obtaining it from I/O, which is itself the
measure of how narrow the guard is. There is no single word for that, and reaching for one is what
produced three wrong labels in a row.
There is now a test that kills that mutant, and the source comment no longer calls it
documentation.
The same reader's proposed instrument closes the general case: make the environment a fixture
parameter instead of a constant. Five forced cells across three axes — controlling terminal
present, fd 0 a tty, fd 1 a tty — with the gate's verdict asserted per cell:
| mutant | killed by which cells |
|---|---|
isatty on both streams |
neither_tty, stdin_only, stdout_only |
isatty on stdin only |
neither_tty, stdout_only |
isatty on stdout only |
neither_tty, stdin_only |
| gate always returns True | no_ctty |
The two mixed cells, where one standard stream is a terminal and the other is not, are the only
thing separating the stdin-only mutant from the stdout-only one. Without them both die to the
same single cell and read as the same result. That is the concrete version of two zeros looking
identical, and my suite had neither mixed cell until this correction.
Population, since it is the whole point of this post: 5 of 131 tests run under a pty, and one
more opens /dev/tty in-process, for six that exercise a terminal-backed path. A pseudo-terminal
is deliberately not evidence of a physical terminal or a person, which is why that phrasing is not
"six tests prove a human."
The count in order, because it is the argument and a reader should be able to check the ordering:
| version | what the gate was | tests opening a terminal |
|---|---|---|
| v0 | a function parameter | 0 |
| v1 |
/dev/tty with r+
|
0 |
| v2 | separate handles | 2 — written as part of this repair |
| v3 | one predicate | 5 under a pty, +1 in-process |
The two terminal tests did not exist while r+ was in the tree. They were written to fix it, and
they are what caught it — restore r+ today and four tests go red, which is the first row of the
mutation table above. Zero is the number that carried v0 and v1, and zero is why both shipped
green.
$ python3 -m pytest -q tests/
131 passed, 25 subtests passed
Population, since it is the whole point of this post: 5 of 131 tests run under a pty, and one
more opens /dev/tty in-process, for six that exercise a terminal-backed path. A pseudo-terminal
is deliberately not evidence of a physical terminal or a person, which is why that phrasing is not
"six tests prove a human."
The count in order, because it is the argument and a reader should be able to check the ordering:
| version | what the gate was | tests opening a terminal |
|---|---|---|
| v0 | a function parameter | 0 |
| v1 |
/dev/tty with r+
|
0 |
| v2 | separate handles | 2 — written as part of this repair |
| v3 | one predicate | 5 under a pty, +1 in-process |
The two terminal tests did not exist while r+ was in the tree. They were written to fix it, and
they are what caught it — restore r+ today and four tests go red, which is the first row of the
mutation table above. Zero is the number that carried v0 and v1, and zero is why both shipped
green.
$ python3 -m pytest -q tests/
131 passed, 25 subtests passed
The honest claim about what this technique is worth
sudoers(5), verbatim from this machine:
requiretty If set, sudo will only run when the user is logged in to a real tty.
When this flag is set, sudo can only be run from a login session and
not via other means such as cron(8) or cgi-bin scripts.
This flag is off by default.
Off by default, and note what the description is about: cron and CGI, not adversaries. Which
is the correct amount of credit to give this. It is not a security boundary:
$ python3 -c 'import sys; print(sys.stdin.isatty())'
False
$ echo "" | script -q /dev/null python3 -c 'import sys; print(sys.stdin.isatty())'
True
script(1) allocates a pseudo-terminal. Code running with my user's permissions and access to
PTY facilities can do that, including the assistant I write most of this code with. Its default
tool calls fail the check — I have watched them fail, and one of my own tests is that wrapper
passing on purpose:
def test_a_pty_wrapper_satisfies_the_gate(self):
"""Stated as a limit, not a defect. This is what the post claims."""
So the claim is narrow: this converts an accidental bypass into a more explicit one. My
current noninteractive automation path no longer gets past it merely by supplying function
arguments. An environment that already controls a PTY can still satisfy the terminal path
without writing anything new, so this is not identity and not authentication. It is worth
having, and it is not the same as being safe.
The second defect, which came back in three days
The ten exports landed in live state because my tests wrote to production paths. I fixed that
on the 24th by redirecting three module constants in the setUp of each class that needed it.
Row 3 of that ledger is the 27th local — 01:31 UTC on the 28th, which is why the timestamp
you read above looks like a different day. A test I wrote during this repair reached
write_receipt() without a redirect and appended to the real ledger. Its refusal_detail is
the old isatty message, which is how I know which version wrote it.
The lesson is not "remember to redirect." Isolation was opt-in, so correctness depended
on every future author choosing to comply, and I was the future author who did not. It is now
opt-out:
@pytest.fixture(autouse=True)
def isolate_agent_state(request):
if request.node.get_closest_marker("live_state"):
yield
return
...redirect RECEIPTS, NONCES, ATTEMPTS to a temp dir...
A test must now declare @pytest.mark.live_state to touch real state, and that declaration
is visible in the test source. Verified load-bearing by flipping autouse=False: four
failures, from tests that exist only to prove the fixture fires.
If a rule in your project is enforced by someone remembering it, it is a request, not a
control. Mine took three days to prove that.
The probe
The tool at the top. Four modes; the one that matters is --redirected, which reproduces the
disagreement:
python3 gate_probe.py # pipeline / CI shape
script -q /dev/null python3 gate_probe.py </dev/null # terminal shape
script -q /dev/null python3 gate_probe.py --redirected # ctty intact, stdio redirected
script -q /dev/null python3 gate_probe.py --detach # no controlling terminal at all
And a cheap smell check, labelled as what it is — a string search, not proof:
grep -rn 'open(["'"]/dev/tty' tests/
A bare grep -rc '/dev/tty' tests/ is worse than useless: it prints one count per file rather
than a single number, and it counts docstrings. On my own suite 9 of its 15 hits are prose
about the terminal. It can only err in the flattering direction, which is the exact failure
mode this post is about. A zero would not prove much either — a test can reach /dev/tty
through production code without the string appearing in tests/ at all.
The evidence that actually settles it is above: a test that exercises the real path without
patching it, plus a mutation showing the suite goes red when that path breaks.
I cannot show you that number for my own broken versions. The directory is a git repo with
history, but the files in question are not in it — git ls-files --error-unmatch returns nothing, and the same for its tests. An earlier version of this
stage2a_executor.py
sentence said "that tree is not under version control," which is wrong in a way a reader could
catch by running git log in it. What I can show is structural and needs no count: v1's approval tests replaced
read_confirmation by name, and v0 had no terminal boundary to exercise at all. A suite that
never opens /dev/tty cannot report anything about a gate that reads from it, whatever its
total.
gate_probe.py, in full
#!/usr/bin/env python3
"""Which question is your human-in-the-loop gate actually asking?
Run it four ways and compare the two verdict columns:
python3 gate_probe.py # a pipeline / CI shape
script -q /dev/null python3 gate_probe.py </dev/null # a terminal shape
script -q /dev/null python3 gate_probe.py --redirected # ctty intact, stdio redirected
script -q /dev/null python3 gate_probe.py --detach # no controlling terminal at all
`isatty` and `/dev/tty` are different predicates, and a gate built on the first while it
reads from the second will disagree with itself. The row that matters is the one where the
two columns differ: that is a process your gate classifies one way and your confirmation
code classifies the other.
Standard library only. Writes no files and mutates no state; it prints to stdout.
"""
import errno
import io
import os
import subprocess
import sys
# Named for what the errnos establish, not for the strongest one in the set. ENOENT means
# /dev/tty does not exist here; ENOTTY means that fd is not a terminal. Neither literally
# asserts "this process has no controlling terminal."
CTTY_UNAVAILABLE = {errno.ENXIO, errno.ENODEV, errno.ENOTTY, errno.ENOENT}
def probe():
"""Return (isatty_verdict, dev_tty_verdict, detail)."""
stdio = sys.stdin.isatty() and sys.stdout.isatty()
try:
out = open("/dev/tty", "w")
inp = open("/dev/tty", "r")
out.close()
inp.close()
return stdio, True, "opened"
except io.UnsupportedOperation as exc:
# Not an unavailable terminal. Reported separately because it subclasses OSError with
# errno None, so an `except OSError` upstream will call this "no human present."
return stdio, False, f"UnsupportedOperation: {exc} (errno={exc.errno}) -- A BUG, NOT AN ABSENCE"
except OSError as exc:
kind = ("controlling terminal unavailable"
if exc.errno in CTTY_UNAVAILABLE else "UNEXPECTED")
name = errno.errorcode.get(exc.errno, "?")
return stdio, False, f"{kind}: errno {exc.errno} ({name})"
def also_show_the_broken_open():
"""The mode that looks correct and is not. A tty is not seekable; text update mode
wants it to be. Reproduced on CPython 3.9.6 and 3.13.9."""
try:
open("/dev/tty", "r+").close()
return 'open("/dev/tty", "r+") -> opened'
except OSError as exc:
return f'open("/dev/tty", "r+") -> {type(exc).__name__}: {exc} (errno={exc.errno})'
def main():
me = os.path.abspath(__file__)
if "--detach" in sys.argv:
# setsid() leaves the session, so the child loses the controlling terminal. Only
# differs from a plain run if the PARENT had one -- run it under script(1) to see it.
r = subprocess.run([sys.executable, me], preexec_fn=os.setsid,
capture_output=True, text=True)
sys.stdout.write(r.stdout + r.stderr)
return 0
if "--redirected" in sys.argv:
# The case that produces the disagreement, and the shape of `pytest` launched from
# a terminal: stdio replaced, session intact. Run this one under script(1).
r = subprocess.run([sys.executable, me], stdin=subprocess.DEVNULL,
capture_output=True, text=True)
sys.stdout.write(r.stdout + r.stderr)
return 0
stdio, ctty, detail = probe()
print(f" isatty(stdin) and isatty(stdout) : {stdio}")
print(f" /dev/tty openable : {ctty}")
print(f" detail : {detail}")
print(f" {also_show_the_broken_open()}")
if stdio != ctty:
print("\n THE TWO PREDICATES DISAGREE for this process.")
print(" A gate on isatty and a confirmation read on /dev/tty will not agree here.")
return 0
if __name__ == "__main__":
raise SystemExit(main())
The question
Not "can a test satisfy your gate" — I had that as the ending for two drafts and my own suite
disproves it. test_a_pty_wrapper_satisfies_the_gate exists on purpose. A test that can drive the
real boundary deliberately is what good integration coverage looks like.
The question is:
Can ordinary automation satisfy your human-in-the-loop gate without crossing the interaction
boundary you intended?
If it can pass merely by supplying the right values into the same function call, then the gate has
established knowledge of those values, not human involvement. v0 established that a caller knew
two strings. It never established anyone read them.
And the follow-up that took me three versions to reach: does any test exercise the real
boundary, or do they all replace it? Mine all replaced it — which is how one suite certified a
gate that accepted everyone and a gate that accepted no one, and reported both as correct.
Mine held for exactly as long as the flag next to it was set to false.
Top comments (24)
The line that should scare anyone shipping an agent with a “human gate” is structural: the same suite certified a gate that accepted everyone (parameters as approval) and a gate that accepted no one (broken
/dev/ttypath), because both versions replaced the interaction boundary in tests.Green did not mean the control worked. Green meant the bypass still compiled.
I would steal two checks from your repair story for any approval / policy / “human present” control:
read_confirmation(or injects the typed ids as arguments), the suite is grading the stub.except, reinstate the staleisattypredicate — if nothing goes red, that line was documentation, not a gate.Narrow claim, same as yours: converting an accidental bypass into an explicit one is worth having, and it is still not identity. The eval contract for the gate is “does ordinary automation pass without crossing the interaction you intended?” — not “did 131 tests pass.”
Your second check is the one that paid off, and the most useful result it gave me was a null.
I mutated a line I had written as a guard, an isinstance check for io.UnsupportedOperation sitting in front of the errno test. Nothing went red. Deleting it fails zero tests, because errno is already None for that exception so the errno check re-raises it anyway. I kept the line and labeled it in the source as documentation rather than a control, specifically so nobody counts it as protection later.
I would add a third check though, because your two would not have caught my last defect.
Count the predicates. My gate asked "are stdin and stdout ttys" and the confirmation read asked "can I open /dev/tty". Two different questions on one control path. No mutation surfaces that, because with no controlling terminal both answers are no, and they only diverge when a controlling terminal exists while stdio is redirected. Which is plain pytest from a terminal. And even there it would not have shown up, because the stale predicate refused first so the real read was never reached. That is how a version that could not open a terminal at all stayed green.
So the third one is just: how many different questions does this gate ask, and is it one function.
The null mutation on the isinstance guard is the result I wish more teams published: green deleted nothing, so the line was never a control. Labeling it as documentation in source is the right move — otherwise the next reviewer counts it as protection.
Your third check is sharper than my two. Mutation alone misses a split control path when every fixture answers both questions the same way. isatty on stdio and “can I open /dev/tty” are not synonyms; they only diverge under a controlling terminal with redirected stdio — which is exactly ordinary pytest from a terminal.
What I would steal as the suite contract for any “human present” / approval gate:
Narrow claim: a control whose questions only agree in the headless CI box can stay green forever and still refuse the real interaction path. The eval is whether automation can pass without crossing the interaction you intended — including the fixture where the old stale predicate would have short-circuited first.
You were right, I did not have it. The test I thought covered that case adapts to whatever environment it lands in rather than forcing one, so in a headless box both answers come back no and it passes without ever exercising the split.
Built it, and building it caught two more of the same thing.
The first version shelled out to script -q /dev/null cmd, which is BSD usage. GNU script takes the command through -c, so that line would not have run the child on Linux, and the test skips when script is absent. A divergence fixture that skips on the reviewer's machine is the vacuous pass one level up. It runs on pty.fork now, which is stdlib and POSIX: opens a pty, forks, setsids, acquires the controlling terminal in the child. Then dup2 of /dev/null onto fds 0 and 1 takes away stdio's tty-ness while the controlling terminal stays.
The second one is the one worth having. Your check needs the fixture to assert it achieved the environment, and I wrote that assert against sys.stdin.isatty and sys.stdout.isatty. Then I deleted my own dup2 calls to see whether the guard would fire. It passed. Under default pytest capture those objects are already replaced and report isatty False, so my guard was being satisfied by the harness rather than by anything the fixture did. os.isatty(0) and os.isatty(1) is the version that actually goes red, because the fds are the only thing that code controls.
Matrix now, under both default capture and -s: clean pass, red when the two predicate gate is restored, red when the redirect is deleted. Before the fd change that last cell was green under default capture, which is the same failure as the article, three levels down.
That matrix under default capture vs
-sis the publishable part — you turned a vague "divergence fixture" ask into a ladder of vacuous greens.Three failure costumes in one thread, same disease:
scriptisn't there / wrong flags — a skip is a quiet pass for gate suites.sys.stdin.isattythat pytest capture already satisfies, so deleting yourdup2still looked green.os.isatty(0)/os.isatty(1)is the right surface because it's the one the code under test can actually own. I'd steal one more contract line for any "human present" / HITL gate: the production predicate and the fixture assert must call the same function. If the suite probes fds while prod still reads stream wrappers (or the reverse), you've rebuilt the disagreement one abstraction up.Narrow claim: a green gate suite that never forces the environment production can land in is still documenting hope, not control.
The same function line is the one I'd steal back, with one catch I found in my own code. The gate and the read do go through one open_controlling_terminal function now, so production can't grow a second predicate. But if that shared function has a bug, the gate and the test just agree with each other about it. So I'd still want the forced fd cells checking the environment they built, not only the gate's answer. Documenting hope instead of control is a good way to put it.
The three versions of this gate are one bug at three levels, and the level that makes it invisible is the fixture, not the code.
v0 accepted everyone because the approval arrived as a parameter, and a parameter is a value any caller supplies. v1 accepted no one because the check asked isatty on two streams while the value came from /dev/tty. Both are the same defect one level apart: the thing being checked is not the thing the value comes from. The repair is not a better proxy, it is a single channel. Require the confirmation to arrive through the same descriptor the gate tests, so that the gate and the read cannot disagree about what they are looking at. A gate over a capability, read from that capability, is the only version of this that cannot certify a different object than the one it then uses.
Then the level that speaks to nomad-link-id. Mutation alone cannot kill this class, and the reason is structural: the mutant is code, the defect is a relation between code and the environment the code runs in. A suite whose fixtures all sit at one environment point cannot distinguish isatty from an openable /dev/tty whatever you mutate, because at that point the two predicates are equal. The instrument is to make the environment a fixture parameter instead of a constant: enumerate the cells (stdin is a tty, stdout is a tty, /dev/tty openable) and assert, per cell, whether the gate refuses. Under that matrix the mutated line dies in the cells where the predicates disagree, and the surviving mutant in the agreeing cells is itself the finding: the guard is load bearing in some environments and decoration in others.
That gives a general form worth writing down. A mutation score is taken over mutants times fixtures. With one fixture environment, everything environment sensitive is unkillable by construction, and zero killed is then a statement about the fixture matrix rather than about the guard. Two zeros look identical on the report.
Your receipt hunt points the same way. Those ten directories prove the runner was entered, because frozen_env is the first statement and the loop is below it. Nothing in the artefact separates entered from started from completed, and the ledger cannot supply the difference either, since its rows were written fifty three minutes later by a different run. The cheapest repair is to make the runner write a start row as its first act, before any side effect, so the three outcomes take three distinguishable shapes. An artefact created before the work is evidence of entry, not of work, and a receipt written in every terminal case still says nothing about a run that never reached one.
I could not re-run the probe here, so what I checked is read rather than executed, and there are two of them. io.UnsupportedOperation is a subclass of both OSError and ValueError, and its errno attribute is None when it is constructed with a message. That is why your isinstance guard is a null mutation, and it is a slightly stronger statement than the one you made: no test written against errno can separate the guarded path from the unguarded one, because errno is None on both sides and the fallback re-raises on the same condition. The documentation label is the right label.
Your mutants times fixtures point cost me a published claim, and a correction is going into the article.
I reported dropping the isinstance guard on io.UnsupportedOperation as a null mutation and concluded the line was documentation rather than a control. That was a statement about my fixture matrix. The cell nothing sampled is not one of your three axes though, it is a fourth one: how the exception gets constructed.
python3 -c "import io; print(io.UnsupportedOperation('msg').errno, io.UnsupportedOperation(6,'x').errno)"
None 6
errno is None only from the message form, and that is what CPython raises for open("/dev/tty","r+"), which is the one shape every test produced. Hand the two arg form ENXIO and the paths separate. Guarded it re-raises. Unguarded it returns "controlling terminal unavailable: errno 6 (ENXIO)", which is the relabel a defect as an absent human bug rebuilt exactly. So the null was an artefact of my fixtures rather than a fact about the line.
What the line is takes two populations, not a better adjective. I tried documentation, then load bearing, then defence in depth, and the third was the same error as the first two because it averaged two answers into one word.
Population A is what the I/O stack raises. Message form, errno None. open("/dev/tty","r+"), a read on a write handle, seek and tell on a pipe, all of them. The errno test already re-raises those, so deleting the line changes nothing on that path, and your last paragraph is right about A.
Population B is the inherited OSError constructor. UnsupportedOperation(ENXIO, "x") carries errno 6, and deleting the line relabels it as an absent terminal. For B the line is the only check.
So redundant against A, sole guard against B, and B only ever arrives by construction. The test that kills the mutant builds the exception instead of obtaining it, which is the measure of how narrow the guard is. Your other read only check holds too, the mro is UnsupportedOperation, OSError, ValueError.
Then I built your instrument, because the above is only the easy half of what you said. Five forced cells over controlling terminal present, fd 0 a tty, fd 1 a tty. os.isatty on the descriptors rather than sys.stdin, because under default pytest capture those objects are already replaced and report False regardless, so a sys level probe measures the harness instead of the cell. I made that exact mistake first and caught it by deleting my own dup2 calls to see if the guard fired. It did not.
Mutants times cells:
isatty on both streams killed by neither_tty, stdin_only, stdout_only
isatty on stdin only killed by neither_tty, stdout_only
isatty on stdout only killed by neither_tty, stdin_only
gate always returns True killed by no_ctty
The two mixed cells are the only thing separating the stdin only mutant from the stdout only one. Without them both die to the same single cell and read as the same result, which is your two zeros exactly. My suite had neither mixed cell before this.
On single channel you are right and I have not fixed it. require_controlling_terminal opens /dev/tty, closes both handles, returns a bool. read_confirmation then opens again. Same path, different descriptor, so the gate certifies a capability and the read acquires it a second time, and the window between them is documented as a race rather than closed. Handing the read the descriptor the gate already proved is a change to both signatures, not a rename.
Start row written before any side effect is the right shape for the entered versus started versus completed problem too. frozen_env creates runtime-tmp as its first statement, which is exactly an artefact created before the work.
Your fourth axis is a real correction, and it lands where my sentence was too broad. I said no test written against errno can separate the guarded path from the unguarded one. That is true of population A and false of B, and B is not a cell the matrix can reach, because it is not an axis of the environment at all: it is a property of the value. A forced-cell design varies where the code stands; the constructor varies what arrives. Five cells cannot see it however many you add, which is a sharper result than the null was.
Verified here before replying, since it is one line: [[UnsupportedOperation("msg").errno]] is None and [[UnsupportedOperation(6,"x").errno]] is 6, and the mro is UnsupportedOperation then OSError then ValueError. One addition to your relabel sentence. The two-argument form also renders its errno into the string: [[str(UnsupportedOperation(6,"x"))]] reads [[[Errno 6] x]]. So the unguarded path does not merely lose the code, it prints a message of the same shape your absent-terminal branch prints. The guard is separating cannot-acquire from absent, and those are two sentences a reader of the log should be able to tell apart.
On redundant against A and sole guard against B, the consequence I would write beside the line: its status is a fact about its callers. B is reachable only by code that constructs the exception, so any wrapper, translation layer or test double that builds one puts you in B, and a refactor that stops building one returns you to A with nothing red. That is why the documentation label is honest at the line, and why the label is worth naming the population rather than the line.
Your double-open is the original defect again, and I would not document it as a race. The gate opens [[/dev/tty]], closes both handles and returns a bool; the read opens it again. So the gate certifies a capability and the read acquires it a second time. The change that closes it is the one you already sized: open once, hand the descriptor into [[read_confirmation]], and let the predicate and the read take the same object as a parameter. Then the window is not narrowed, it is gone, because there is no second acquisition that can fail. The signature change is the whole fix, which is your v0 observation running backwards: there the defect was a value any caller could supply, here the fix is a value no caller can supply.
On the start row, I would put it as the first statement of [[run_frozen_commands]], before [[frozen_env]], and give it a state field that the completion and failure paths rewrite. Then entered, started and completed are three values of one row rather than three artefacts that each mean a different thing. [[frozen_env]] creating [[runtime-tmp]] as its first statement is exactly the counterexample that motivates it.
Your matrix row is the part I would keep in a methods note. The two mixed cells are the only thing separating the stdin-only mutant from the stdout-only one; without them both die to one cell and read as the same result. A cell in which two mutants die for the same reason cannot tell you which one you are looking at, which is the null-mutation failure one level up.
Both of these are still open, so I'll just say it straight. The gate and the read share one open function now, but the gate still opens /dev/tty, closes it, and then the read opens it all over again. So your double open is there today. Sharing the helper didn't fix it, handing the handles through would.
Same with the start row. runtime-tmp is still the only sign the runner was even entered. The one thing I'd do differently is keep that row out of the receipt ledger. The ledger's append only, and a row that gets rewritten from entered to completed would wipe out its own history.
The str() thing is a really good catch too. If both paths end up printing [Errno 6], putting the exception type next to the errno is what lets someone reading the log tell can't acquire apart from not there.
"Does any test exercise the real boundary, or do they all replace it?" is a question I'm going to start asking about agent gates too. The same trap shows up there: every test mocks the check, so a gate that approves everything and one that blocks everything look identical to the suite. The cheapest guard I know is a pair of tests per gate, one input that must pass and one that must be refused, both going through the real boundary. If either can't be written, the gate isn't really there. Has this changed how you look at the other gates in that codebase?
Yeah, it has. I went through the policy gate sitting right next to it with your pair in mind and found a gap straight away. The refusal tests check the verdict and the receipt, but none of them checks that a refusal left zero attempt directories behind. The must pass side can run against an isolated policy fixture without touching the real one, as long as it checks the thing actually happened and not just that it made it to the next gate.
That's a good find, and it's the kind tests rarely look for: a refusal is only a refusal if nothing happened. Checking the verdict and the receipt proves the gate said no; checking for zero attempt directories proves the system listened. I'd go one step further and make "no side effects" a reusable assertion for every must-refuse test, so the next gate gets it by default. Was the gap real, or did the directories just never get created in the first place?
The fixture point is already in the thread, so the thing I would pull out instead is ten attempt directories against zero receipts. Two records of the same run disagreeing about whether it happened, and neither of them lying. The directories say the runner was entered. The ledger says nothing was authorised. That is a real state rather than a bug in either writer, and it is only visible because you read both.
The generalisable part of the isatty story is smaller than the terminal detail makes it look. Your gate read one handle and your confirmation read another. They answer different questions, and they agree often enough that the disagreement arrives as a three-day mystery instead of a failure. Same shape in payments any time the authorisation check and the movement of money read different systems. The rule that survives is that the predicate which decides has to come off the same object the effect will use, not off a second one that usually matches.
On the tests: a suite that cannot separate accepts-everyone from accepts-nobody is a suite with one arm. Neither version fails an assertion on the happy path, because the happy path is the only path with an expected outcome written down. The assertion that separates them is not about either run. It is about the delta. A refusal leaves zero attempt directories, an approval leaves exactly one, and the test asserts the difference across two runs that vary only in the human answer. Your ls output would have failed that on the first run.
Where I would stop short of you is the flag flip. Turning execution on is the only reason the post exists, and it also means every number here was produced in a configuration the policy forbids. I would want the identical probe run with the flag off, asserting zero attempts and zero receipts, because that is the state the system is meant to be in and the one nobody ever tests.
The delta assertion is the one I'm missing. Nothing in the suite checks that a refusal leaves the attempts directory empty. The policy refusal happens to return before attempt_root.mkdir, but that's code order, not something a test holds in place. Zero new directories on refusal, exactly one on approval, two runs that differ only in the answer is the test that would.
Where I'd push back is zero receipts. The executor writes a receipt in every terminal case on purpose, because a refusal that leaves no record isn't auditable. So for me the flag off expectation is zero attempts and one refusal receipt.
One boundary you can exercise for real: in the agent harness I use, a denied tool call comes back to the model as feedback, not an error. So the refusal path is observable end to end. Deny once and watch whether it adjusts or retries the same call verbatim, which just re-prompts the human.
That's a good one. Whether it changes approach or just fires the same call again is exactly what I'd want to see, because a verbatim retry turns the human gate into a nag loop. This executor doesn't call a model yet though, so right now all I can test is the refusal and what it leaves behind. The adapt or retry part has to wait until there's actually a loop.
The isatty-versus-/dev/tty split is a sharp catch. I lost an afternoon to almost this exact thing: a check that read stdout.isatty() behaved one way in my terminal and the opposite under pytest's default capture, and no assertion I had written could see the difference because both green states looked identical. The probe-that-writes-no-files approach is the honest way to pin it down. Did you end up gating on the /dev/tty open succeeding, or on something else entirely?
Yeah, on the open succeeding, just a bit narrower. It opens /dev/tty as two separate handles, one to read and one to write, and the gate only passes if both open. If it fails with ENXIO, ENODEV, ENOTTY or ENOENT it refuses and writes that errno into the receipt. Anything else gets raised, so a bug can't come out looking like nobody's there. Worth saying it only proves a terminal is available, not that a person is sitting at it.
And the pytest capture thing was my afternoon too. Both greens look exactly the same.
The part about tests passing while the actual security control was broken is pretty eye-opening. Testing the real interaction boundary instead of mocking it seems like the big takeaway here.
The part that still gets me is what was actually protecting it. Not the typed confirmation, just a boolean in a config file I'd left set to false. Two separate controls, and a green test on one tells you nothing about the other.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.