A replay that can still rewrite fixtures is not a merge signal. It is a write that happened to exit zero.
Freeze the fixture digest first. Then replay the touched tests on a tree that cannot update those files. Unlock merge only when the digest before and after the replay is the same, and the flake window stays inside a budget you chose. The commands below are an unexecuted proposal. Adapt the paths. Do not treat them as a transcript from a live pipeline.
What a single green job hides
You already know a test can pass for the wrong reason. The annoying case is narrower than a flaky network. The job edits a fixture, the assertion matches the new bytes, and the check turns green.
That pattern shows up when an agent drafts tests, and it shows up when a person updates snapshots in a hurry. The exit code does not record which one happened.
Split the permissions. One step may classify the diff and record a digest. A later step may run tests. The later step must not be the one that updates fixtures. If those two steps share a writable worktree, you did not split anything.
Three artifacts, one rule
You only need three files to make the split reviewable.
- A path list of fixture roots, maintained by people, not inferred from filenames.
- A classify hook that fails closed when those paths change without an explicit label.
- A replay script that checks the digest, drops write access as far as the runner allows, reruns the touched tests, and writes a receipt.
The rule is small. No receipt, no merge. A local green, an agent summary, or a log from a side machine does not replace the receipt on the SHA you are merging.
Decide before you add a check
Use a table so the hook stays dumb. Ambiguous diffs fail. They do not get a score.
| Diff you see | Classify | Replay | Merge |
|---|---|---|---|
| App code only | Pass, no digest | Skip | Your existing required checks |
| Tests changed, fixture roots untouched | Pass, list test paths | Rerun those tests | Allow only inside the flake budget |
Fixture root changed, no fixture-update label |
Fail | Do not start | Block |
| Fixture root changed, label present, digest recorded | Pass | Rerun, then recheck digest | Allow only if digest is unchanged and budget holds |
| Gate workflow and fixtures in the same PR | Fail closed | Do not start | Split the PR |
Read the last row as a review rule. A pull request that edits the gate and the evidence together is not something this receipt can vouch for.
Step 1: Name the roots
Put the roots in a file the hook can read with no extra parser.
# ci/fixture-roots.txt
tests/fixtures/
tests/http_fixtures/
src/components/__snapshots__/
You own every line. If a team stores fixtures beside source under another folder name, add that folder. Do not let a comment in the diff redefine the list.
Keep globs out until the hook has a real matcher. A prefix is easier to explain in review, and easier to fail closed when a path surprises you.
Step 2: Fail closed on unlabeled fixture diffs
The hook does not decide whether the new fixture is true. It decides whether a human allowed the write, then it records a digest.
#!/usr/bin/env bash
# proposal, unexecuted: ci/hooks/check-fixture-label.sh
set -euo pipefail
roots_file="${1:-ci/fixture-roots.txt}"
label="${FIXTURE_UPDATE_LABEL:-}"
base="${BASE_SHA:-origin/main}"
mapfile -t roots < <(grep -Ev '^\s*(#|$)' "$roots_file")
changed=$(git diff --name-only "${base}...HEAD")
hits=()
while IFS= read -r path; do
[ -z "$path" ] && continue
for root in "${roots[@]}"; do
case "$path" in
"$root"*) hits+=("$path") ;;
esac
done
done <<< "$changed"
mkdir -p ci/receipts
if [ "${#hits[@]}" -eq 0 ]; then
echo "classify=no-fixture" | tee ci/receipts/classify.txt
exit 0
fi
if [ "$label" != "fixture-update" ]; then
printf 'unlabeled fixture path:\n' | tee ci/receipts/classify.txt
printf '%s\n' "${hits[@]}" | tee -a ci/receipts/classify.txt
exit 1
fi
# sha256sum splits on spaces; keep fixture paths space-free or switch to a NUL-safe digest
sha256sum "${hits[@]}" | sort > ci/receipts/fixture-digest.txt
echo "classify=frozen count=${#hits[@]}" | tee ci/receipts/classify.txt
Try the failure path before you try the success path:
FIXTURE_UPDATE_LABEL="" BASE_SHA=origin/main bash ci/hooks/check-fixture-label.sh
You want a non-zero exit when a listed root appears in the diff. Then repeat with FIXTURE_UPDATE_LABEL=fixture-update and confirm ci/receipts/fixture-digest.txt lists only those paths. If the digest lists a source file, your roots are too wide. Fix the list before you write the replay.
Step 3: Replay without a writable fixture tree
chmod a-w is not a boundary. The same user can turn the bit back on and write the file. Treat a mode change as a tripwire only.
Prefer a stronger lock when the runner allows it. Copy the fixture tree, bind-mount it read-only over the working paths, and run the tests as a user who cannot mount or change mode. If you cannot get a read-only mount, say so in the receipt and do not pretend the lock held.
#!/usr/bin/env bash
# proposal, unexecuted: ci/replay-locked.sh
set -euo pipefail
digest_file="${1:-ci/receipts/fixture-digest.txt}"
test_cmd="${2:-pytest -q tests/test_order_replay.py}"
budget="${FLAKE_BUDGET:-3}"
lock_mode="${LOCK_MODE:-unset}"
[ -f "$digest_file" ] || { echo "missing digest; classify did not freeze fixtures"; exit 1; }
sha256sum -c "$digest_file"
before_ok=$?
if [ "$before_ok" -ne 0 ]; then
echo "digest mismatch before replay"
exit 1
fi
if [ "$lock_mode" != "ro-mount" ]; then
echo "lock_mode=${lock_mode}; refusing to call this a write-lock" | tee ci/receipts/replay.txt
exit 1
fi
fails=0
i=1
while [ "$i" -le "$budget" ]; do
if ! bash -lc "$test_cmd"; then
fails=$((fails + 1))
fi
i=$((i + 1))
done
if ! sha256sum -c "$digest_file"; then
echo "digest moved during replay" | tee ci/receipts/replay.txt
exit 1
fi
if [ "$fails" -ne 0 ]; then
echo "flake window ${fails}/${budget}" | tee ci/receipts/replay.txt
exit 1
fi
printf 'replay_ok fails=%s budget=%s lock=%s\n' "$fails" "$budget" "$lock_mode" \
| tee ci/receipts/replay.txt
The script exits 1 unless you deliberately pass LOCK_MODE=ro-mount after you have actually mounted the tree read-only. That default is the point. A green receipt must not appear just because the variable was forgotten.
A budget of three is a filter for tests that fail every other run. It will not catch a rare flake. Write that limit into the job summary so the number is not mistaken for a confidence interval.
Paths with spaces will break sha256sum field splitting. Ban spaces in fixture paths, or replace the digest with a NUL-safe tool you already trust. Do not discover that in the merge queue.
Step 4: Require both receipts on the same SHA
Wire classify and replay as separate jobs. The merge job only checks files. It does not rerun tests, and it does not regenerate fixtures.
# proposal, unexecuted — pin the action SHAs your org already reviews
name: fixture-replay
on:
pull_request:
jobs:
classify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout
with:
fetch-depth: 0
- id: lab
run: |
val=""
while read -r name; do
if [ "$name" = "fixture-update" ]; then val="fixture-update"; fi
done < <(jq -r '.pull_request.labels[]?.name' "$GITHUB_EVENT_PATH")
echo "value=${val}" >> "$GITHUB_OUTPUT"
- env:
FIXTURE_UPDATE_LABEL: ${{ steps.lab.outputs.value }}
run: bash ci/hooks/check-fixture-label.sh
- uses: actions/upload-artifact
with:
name: fixture-receipts
path: ci/receipts/
replay:
needs: classify
runs-on: ubuntu-latest
steps:
- uses: actions/checkout
- uses: actions/download-artifact
with:
name: fixture-receipts
path: ci/receipts/
- name: Mount fixtures read-only before this step
run: |
echo "Replace this step with the ro mount your runner actually supports."
echo "Do not export LOCK_MODE=ro-mount until that mount is in place."
exit 1
- run: LOCK_MODE=ro-mount FLAKE_BUDGET=3 bash ci/replay-locked.sh
- uses: actions/upload-artifact
with:
name: replay-receipt
path: ci/receipts/replay.txt
merge-receipt:
needs: replay
runs-on: ubuntu-latest
steps:
- uses: actions/download-artifact
with:
name: replay-receipt
path: ci/receipts/
- run: |
grep -q '^replay_ok ' ci/receipts/replay.txt
grep -q 'lock=ro-mount' ci/receipts/replay.txt
Two details are easy to get wrong. merge-receipt must download the replay artifact. A job that only tests a path it never fetched will fail, or worse, pass if a stale file was baked into the image. The action names above are placeholders. Pin the SHAs your org already reviews. Do not copy a floating tag into a required check.
Watch merge-receipt fail on an unlabeled fixture edit before you depend on it. A check you have never seen fail is a decoration, not a gate.
Draft the failing test somewhere cheap, then throw that green away
You still need a scratch loop for the first red test. That loop should not be the merge runner.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode's free model access can draft the failing test and a first pass at fixture-roots.txt. The free server option is a fit for one rehearsal of the replay script before the change spends shared CI minutes. This article does not name a model, a token quota, a machine size, or a retention period. Those are product terms, and they move. Read the current terms before you plan around them.
Keep that session in draft territory. Copy the test command into the repo. Let the required job produce the receipt on the pull request SHA. A green log from the free server is a rehearsal. It is not the merge signal, even when the same script printed replay_ok.
If you want that rehearsal before the first push, use the free server to confirm the test fails for the reason you think, then land the hook in a separate pull request that does not touch fixtures.
Who should not use this
Skip the pattern when you cannot name fixture roots without guessing. A hook with a vague list becomes a bypass target in a week.
Skip it when fixtures live in object storage or a shared database. A read-only mount on the runner does not freeze those bytes, and a replay may write into a shared environment.
Skip it when the suite is not safe to run three times. Billing jobs, email jobs, and tests that mutate a shared tenant will turn the flake window into an incident.
Skip it when policy forbids a third-party drafting server from seeing fixture contents. Use an internal runner for the rehearsal. The merge rule does not require the free server. The free server is only a place to draft.
Also skip it if fixture-update is applied by a bot. A label a workflow can add is not a human allow. Remove that automation, or drop the label check and review fixture diffs by hand.
Limitations you should write next to the check
The digest proves the listed files did not change during the replay. It does not prove the fixture was correct, complete, or free of secrets.
The flake budget proves the command stayed green for N runs on one runner. It does not prove the test is stable on another architecture, at another hour, or under load.
jq on $GITHUB_EVENT_PATH assumes a pull_request event with a labels array. A workflow_dispatch run will not have that shape. Do not reuse this workflow file unchanged for every event.
The classify hook trusts BASE_SHA. If that SHA is attacker-controlled, the diff can be made empty. Set the base from your protected branch in the workflow, not from a pull request input.
What to land first
Land ci/fixture-roots.txt and the classify hook in a pull request that does not change fixtures. Watch it fail on a deliberate unlabeled edit. Add the replay job only after that failure is visible. Require merge-receipt last.
Until both receipts name the SHA under review, the path is not green. Do not merge on the rehearsal.
Top comments (0)