Daniel Nguyen
← All work

A coverage check in CI that cannot be diluted

I replaced a global coverage percentage with a CI check that fails a pull request when any changed line or branch is untested.

TTMI
Every changed branchUntested lines and branches fail the PR
Area
Retail ERP platform
When
2025–2026
My role
Designed and implemented
  • Python
  • pytest
  • coverage.py
  • GitHub Actions
  • Git

How it fits together

A coverage check in CI that cannot be dilutedThe coverage check joins the pull request's changed statements with line and branch data and fails on any untested change.DeveloperpersonPull requestserviceTest suiteworkerLine + branch reportdata storeDiff analyzerserviceStatement wideningserviceCoverage checkcheck or alarmPR statuspersonopens / pushesrun with branch dataproducesdiff vs merge basechanged linescovered lines+brancheschanged statementspass / fail + gaps
The coverage check joins the pull request's changed statements with line and branch data and fails on any untested change.
  • Person
  • Service
  • Worker
  • Data store
  • Check or alarm

The problem

Our retail ERP is a large Python/Django codebase that grows quickly and has many contributors. Test coverage was enforced as one global percentage with broad exclusions. As long as the number stayed above a threshold, the pull request passed.

The flaw is that a global percentage measures the whole repository and ignores the change under review. If the codebase already has thousands of tested lines, a pull request can add a few hundred untested ones and the number barely moves. The existing tests average the new gap away. The check passes in exactly the case where it should fail, and it gets weaker as the codebase grows.

I designed and implemented a coverage check that looks only at what a pull request changes, and wired it into the pull-request workflow.

What I did

Check the diff instead of the total

The coverage check computes which lines a pull request changed by diffing against the merge base, then requires every one of those lines to be covered by tests. Old code is not judged, so nobody has to fix years of history before shipping a small change. The trade-off is real friction: a large pull request now needs full tests for everything it touches. A higher global threshold would have been easier to adopt on the first day, but it can always be gamed by volume.

Require branch data, and refuse reports without it

Line coverage alone misses a common case. An if statement runs its line whenever it is evaluated, even if only one side was ever taken. So the check requires branch measurement and fails outright if the coverage report was produced without it. That closes the gap where a condition looks tested because its line executed once. The cost is slightly slower test runs and more failures to explain at first.

Widen changed lines to whole statements

A changed line is often part of a longer statement that wraps across several lines. If only the first line counts as changed, an untested continuation can hide below it. The check uses Python’s syntax tree to expand each changed line to the full statement it belongs to. This adds some complexity to the check itself, but it removes a quiet way around it.

No escape hatch for application code

The only thing excluded from the check is the check itself. If someone wants to exclude application code, that exclusion has to appear in the diff where reviewers can see and question it. This is stricter than the old setup with its broad exclusions, and it occasionally means writing a test for code that feels too simple to test.

Error output that tells you what to write

A strict check with vague output makes people fight it. On failure, the check lists each uncovered location and each branch destination that was never reached, so the author knows exactly which test is missing. Getting this output right took about as long as the core logic.

Result

The check runs in CI on pull requests that touch the largest application area of the codebase. New code in that area can no longer lower the effective coverage of the change under review, because untested changed lines and branches fail the pull request.

Review conversations also changed. Instead of asking whether coverage is still above some number, reviewers and authors talk about which branch of this specific change has no test.

What I’d do differently

If I did it again, I would roll it out in a warn-only mode for a couple of weeks before making it blocking, so the team could see the output and give feedback before it stopped their merges. I would also extend it to the rest of the codebase sooner, since an area outside the check is exactly where coverage can still slip.