My reading of the unlazy README is that its real product is a refusal to accept "done" without evidence, and that the parts worth reading closely are the statements about what its own checker cannot vouch for. Much of the README covers the mechanics behind that: an acceptance ledger, runnable gates, and re-verifying work that comes back.
A ledger written before the work
The README's tagline sets the order of operations: "Write the acceptance ledger first. Execute reviewed checks. Reverify returned work. Report only what the evidence supports." For a solo task, that ledger is a template copied to GATES.md. In the README's example, each runnable gate has a CHECK: line with shell code, an EXPECT: line, and an EVIDENCE: line set to pending, and one gate adds a CWD: line.
The pass condition has two parts. According to the README, a runnable gate passes only when its process exits 0 and EXPECT: matches combined output, and output past the 1 MiB limit is never truncated into success. Automatic evidence begins with a versioned SHA-256 digest of the parsed CHECK: and EXPECT: lines and the raw CWD: line, and a checked runnable gate whose evidence does not match its definition is stale and unmet. The parser also rejects ledgers with zero gates, duplicate ids, incomplete runnable gates, or invalid expectations, along with abandonment entries that lack a reason or name an unknown gate id.
Abandonment gets its own outcome. For a valid abandonment, the checker exits 1 with HANDOFF REQUIRED, and the README describes that as a terminal handoff rather than success.
Review before execution
The CLI is staged. The README says --status is the only mode that is always non-executing. A normal run against a new oracle with no exact approval record prints the resolved command, expectation, working directory, shell, and PATH without executing it. The README warns against treating that as a permanent dry run, because once the exact oracle is approved, normal mode can execute it. You approve with --approve after reading every command and called script, and --reverify re-runs all runnable gates, including ones already marked complete.
Approval records live under ~/.unlazy/approved by default. Each record is bound to the absolute ledger and gate, the exact CHECK: and EXPECT:, the resolved CWD: and shell, the timeout, output and regex limits, regex worker limits, the platform, and the full inherited PATH. Editing any bound input requires approval again.
Limits of the checker
This is the section I would read twice before adopting the tool. The README states that the checker "can prove only the command oracle you declare" and cannot infer that an English gate title and arbitrary shell code mean the same thing. It says the definition binding "detects structural drift, not ledger tampering," since anyone who can edit a ledger can forge canonical-looking evidence. It says "Approval is consent, not a sandbox": approval does not hash called scripts, fixtures, or dependencies, and checks run with ambient filesystem, environment, credential, and network access.
The parallel work section uses the same framing. Ready leaves may run together only after each declares disjoint OWNS: paths and claims them, but the README calls lease matching "a coordination guard, not write isolation" and suggests separate worktrees for colliding worktree-local output. In its shell and PATH notes, the README adds that on Windows a checker launched from Git Bash can see Unix-like tools that the same checker launched from PowerShell does not, and it counts a shell or PATH mismatch during parent re-verification as a failed verification.
So the quality of a gate is on you. The README lists habits of good gates, such as printing a success-only marker after all assertions pass and measuring supplied figures instead of copying them into EXPECT:. An advisory, non-executing scripts/gate-lint.mjs flags mechanically weak ledger patterns, and --strict makes warnings fail.
Orchestration and the Stop hook
For work that needs fresh contexts, the README describes a scoped pipeline under .unlazy/<scope>/ with a PLAN.md, a GATES.md, and leaf and node gate files. Leaves use WAITING, READY, IN-FLIGHT, VERIFIED, or ABANDONED states. Gate checks stay sequential by default, with --jobs <N> as an opt-in limit for independent checks. Running gate-check.mjs --scope <id> prints ALL MET only when every gate is met and every launch wave is complete.
An optional Claude Code Stop hook returns a decision: "block" response while gates remain unmet or launch waves remain incomplete. It does not execute checks, and the README says its progress guard releases after six consecutive blocks without semantic gate or dispatch progress. The README asks that it be installed only with the user's consent; the default installation writes .claude/settings.local.json.
Getting it
Install with npx skills add Leonxlnx/unlazy, or clone into ~/.claude/skills/unlazy for Claude Code or ~/.codex/skills/unlazy for Codex CLI. The checker and optional hook require Node 16 or newer and use no third-party runtime packages. On versioning, the README says the current source targets 2.1.0 and "is not identified here as a tagged GitHub release," and it recommends pinning an exact commit when you need an immutable installation.
GitHub: https://github.com/Leonxlnx/unlazy
Curated by Agent Palisade — practical AI for small and mid-sized businesses.
Top comments (2)
The distinction between structural drift and write isolation is the sharpest point in that breakdown. A hashed command string guarantees the command text did not change across turns, but it does not stop an agent from rewriting a test fixture or adding a mock that turns a failing assertion into an exit code 0. In practice, the verification oracle has to run from a read-only harness worktree where the agent has no write access, leaving only the source tree mutable.
The "limits of the checker" section is the part I'd want every tool like this to have, because the gap it names is exactly where I got burned. A command oracle can prove the command ran and matched; it can't prove the English title meant what you think. My version: an agent posted on a schedule for three days, and its gate was "the API returned 200 and a post ID." Every gate passed. The posts reached 2 to 8 people. The gate was true and the title ("post went out") was, in any sense I cared about, false. Writing the acceptance ledger first would have forced me to write EXPECT: impressions > N, and I'd have discovered on day one that I couldn't check it from the server. Which is the useful failure.