Daniel Nguyen
← All work

Agent playbook: git gates that keep AI roles honest

An open-source rule and skill pack for AI coding agents. Git-level gates stop a tester, implementer or reviewer agent from changing files outside its role.

Open source
2,325 tests, 3 OSes165 commits, 17 tagged releases
Area
Developer tooling for AI agents
When
2026
My role
Author and maintainer
Source code
GitHub
  • Bash
  • PowerShell
  • Git
  • GitHub Actions

How it fits together

Agent playbook: git gates that keep AI roles honestThree agent roles hand work through a git-level role gate; only gated commits reach CI, and releases are cut only after main is green.OrchestratorpersonTester agentworkerImplementer agentworkerReviewer agentworkerRole gatecheck or alarmGit historydata storeEvidence logdata storeCI: Linux/macOS/Wincheck or alarmTagged releaseexternalspec, no impl hintsspec + RED commitspec + diff + evidenceonly test pathsonly code pathsno changescommit if gate passesverdict + test outputpushtag after green main
Three agent roles hand work through a git-level role gate; only gated commits reach CI, and releases are cut only after main is green.
  • Person
  • Worker
  • Check or alarm
  • Data store
  • External

Try the role gate

The role gate classifies every file changed since a base commit as test, code or test infrastructure and fails if the current role touched a kind it may not change.

You are the
Files you changed
Gate says

    A simplified model of the real script.Simplified from the real script. Kind 'source' is the script's 'code' (any path that is not a test or infra path, so README.md counts as code). 'handoff' is not a kind in the script: files under .agents/handoff/ are skipped before classification. 'infra' only exists when a project sets its opt-in test-infrastructure pattern; the demo assumes one that matches ci/, and by default ci/run-tests.sh would classify as code. Infra is matched before the test pattern. The real gate reads changed files from git (committed since a base ref, staged, unstaged and untracked, renames as delete plus add), prints one VIOLATION line per offending file and a summary, and exits 1 if any exist. Exit 3 also covers environment errors (not a git repo, unknown base ref); an empty or invalid test or infra pattern exits 3 instead of matching nothing or everything (fail closed). Usage errors exit 2. The real test pattern also covers spec/e2e/fixtures/__mocks__ folders, .spec./_spec. files, conftest.py, .feature files, *Test(s).java/kt/cs/swift, and jest/vitest/playwright/cypress/karma/mocha/phpunit config. The separate Lite size check (3 files, 100 lines) is not modelled.

    The problem

    AI coding agents write code fast and cut corners in predictable ways. They report “done” without a fresh test run. They change files nobody asked them to touch. The worst one is quiet: when the same agent writes the code and the tests, the tests tend to agree with the code’s mistakes, so everything goes green and the bug ships.

    The usual fix is a longer instruction file. That helps, but prose is a request. A model can skip it, and the gap shows up only in review. I wanted the important rules to be checked by something that does not depend on the model’s mood: git, exit codes and CI.

    agent-playbook installs one short always-on rule block and five skills into three agent harnesses, including Claude Code and OpenCode. The core is test-driven development with three separated roles. A tester agent writes failing tests, an implementer agent makes them pass, and a reviewer agent re-runs everything read-only. I designed the rules and the gates; coding agents did much of the work under those same gates.

    What I did

    Enforce roles with git instead of trust

    The role gate looks at every path changed since a base commit, including staged, unstaged and untracked files, and classifies each one as test, code or test infrastructure using a path pattern. Each role has an allowed set. The tester may only touch tests, the implementer may only touch code, the reviewer may change nothing. One file out of bounds and the gate exits with code 1, naming the file. Hand-off notes and evidence logs live in a folder the gate always allows.

    The trade-off: the check works by path. It cannot catch an implementer that special-cases test inputs inside production code. The reviewer role exists for that, and the README says so plainly.

    Fail closed when the configuration is wrong

    Projects can override the test-path pattern. If the override is empty or invalid, a naive check would match nothing, or everything, and every file would quietly pass. The gate refuses to run instead and exits with a separate configuration error code. Snapshot files also count as tests: an implementer who rewrites a snapshot makes a failing test pass without changing behaviour. The cost is that a project whose snapshot files are not tests must set its own pattern.

    Treat portability as a requirement

    The scripts are plain Bash with no dependencies. CI runs the full suite on Ubuntu, on macOS (Bash 3.2 and BSD tools) and on Windows under Git Bash. Windows starts processes slowly, so its run is split into three shards, assigned greedily from measured per-suite times. Old Bash means giving up newer shell features, and the test runner works around the missing ones.

    Release only what main already proved

    A version tag does not re-run tests. The release job waits for the main-branch run of the same commit, requires it to succeed, checks that the tag, the version file and the changelog agree, and never overwrites an existing release. Releases are a little slower; a release can never come from untested code.

    Pick the process weight by risk

    Not every change needs three agents. The playbook has Full, Lite and Exempt weights. A size check counts production files and changed lines (default limits: 3 files, 100 lines) and pushes an oversized “small” change up to Full.

    Result

    The repo has 165 commits and 17 tagged releases so far, up to v0.16.0. The latest main run is green on all three operating systems, and its Ubuntu log shows 2,325 passing tests across 21 suites with none failing. The installer edits only its own managed blocks, keeps backups, and can roll back to an earlier tag. The project is still marked alpha.

    What I’d do differently

    The gate only knows paths, so I would move the mutation spot-check from the reviewer’s checklist into CI. I would also run the behaviour evals across more models and harness versions; right now the dated results cover one harness release. Prose rules still depend on the model following them, and I would keep moving more of them into mechanical checks.