The problem
AI coding agents write code fast and cut corners in predictable ways. They report “done” without a fresh test run. They change files nobody asked them to touch. The worst one is quiet: when the same agent writes the code and the tests, the tests tend to agree with the code’s mistakes, so everything goes green and the bug ships.
The usual fix is a longer instruction file. That helps, but prose is a request. A model can skip it, and the gap shows up only in review. I wanted the important rules to be checked by something that does not depend on the model’s mood: git, exit codes and CI.
agent-playbook installs one short always-on rule block and five skills into three agent harnesses, including Claude Code and OpenCode. The core is test-driven development with three separated roles. A tester agent writes failing tests, an implementer agent makes them pass, and a reviewer agent re-runs everything read-only. I designed the rules and the gates; coding agents did much of the work under those same gates.
What I did
Enforce roles with git instead of trust
The role gate looks at every path changed since a base commit, including staged, unstaged and untracked files, and classifies each one as test, code or test infrastructure using a path pattern. Each role has an allowed set. The tester may only touch tests, the implementer may only touch code, the reviewer may change nothing. One file out of bounds and the gate exits with code 1, naming the file. Hand-off notes and evidence logs live in a folder the gate always allows.
The trade-off: the check works by path. It cannot catch an implementer that special-cases test inputs inside production code. The reviewer role exists for that, and the README says so plainly.
Fail closed when the configuration is wrong
Projects can override the test-path pattern. If the override is empty or invalid, a naive check would match nothing, or everything, and every file would quietly pass. The gate refuses to run instead and exits with a separate configuration error code. Snapshot files also count as tests: an implementer who rewrites a snapshot makes a failing test pass without changing behaviour. The cost is that a project whose snapshot files are not tests must set its own pattern.
Treat portability as a requirement
The scripts are plain Bash with no dependencies. CI runs the full suite on Ubuntu, on macOS (Bash 3.2 and BSD tools) and on Windows under Git Bash. Windows starts processes slowly, so its run is split into three shards, assigned greedily from measured per-suite times. Old Bash means giving up newer shell features, and the test runner works around the missing ones.
Release only what main already proved
A version tag does not re-run tests. The release job waits for the main-branch run of the same commit, requires it to succeed, checks that the tag, the version file and the changelog agree, and never overwrites an existing release. Releases are a little slower; a release can never come from untested code.
Pick the process weight by risk
Not every change needs three agents. The playbook has Full, Lite and Exempt weights. A size check counts production files and changed lines (default limits: 3 files, 100 lines) and pushes an oversized “small” change up to Full.
Result
The repo has 165 commits and 17 tagged releases so far, up to v0.16.0. The latest main run is green on all three operating systems, and its Ubuntu log shows 2,325 passing tests across 21 suites with none failing. The installer edits only its own managed blocks, keeps backups, and can roll back to an earlier tag. The project is still marked alpha.
What I’d do differently
The gate only knows paths, so I would move the mutation spot-check from the reviewer’s checklist into CI. I would also run the behaviour evals across more models and harness versions; right now the dated results cover one harness release. Prose rules still depend on the model following them, and I would keep moving more of them into mechanical checks.