The problem
AI coding sessions usually run every task on one model at one effort level. A rename and a cross-module refactor get the same expensive model. That over-spends on routine edits, under-powers risky ones, and leaves no record of which model did the work or whether it was correct.
I wanted a router that picks the cheapest model tier that should still do the job, escalates when the risk is real, and writes down what happened in a form I can check later. The hard part is honesty: a router that guesses its own success rate would have me tuning it on false numbers.
I built it as one router with two adapters. The difficulty-routing adapter drives one coding-agent CLI and routes by task difficulty. The role-routing adapter drives a second coding-agent CLI and routes by role in a test-first workflow. Both share the Python routing scripts, a PowerShell wrapper and one telemetry format.
What I did
Route on risk signals computed without a model
Before any model runs, a deterministic preflight puts the request in a category, scores which repository files look relevant, and derives signals such as complexity, uncertainty, blast radius and whether the work crosses components. The difficulty-routing adapter sends work to a free executor by default. It escalates to a premium executor when a signal is high, the work is cross-component, a prior attempt failed, or I ask for it. A planner model is called only in those cases, with a small, secret-redacted context packet. A fixed share of medium tasks goes premium on purpose, as a comparison group. The trade-off: keyword-based preflight is crude and can misjudge a task. I accepted that because it is cheap and repeatable.
Cap the loop in code
For executable repo work, the adapter runs the executor, then a verifier, then a planner verdict: done, continue, or ask me. An earlier version stated its call limit only in the instructions, and the history shows the limit was exceeded. I moved it into code as a hard cap on iterations and delegated calls. When the budget runs out, the router asks me instead of trying again. The verifier must report real exit codes; a check it could not run counts as unverifiable, never as a pass. Sometimes work stops a little early; I prefer that to an unbounded bill.
Keep three verdicts separate
Each run records three things that are never inferred from each other: did the task succeed, did the routing follow policy, and was the telemetry written. A route counts as observed only when it comes from runtime metadata or an execution receipt. A missing receipt is labelled unverified, and a contradiction is labelled noncompliant. Telemetry never stores prompts, command text, tool output or credentials. If the sync permission is missing, records wait in a local queue and are deduplicated later. The cost: many records stay “unobserved”.
Measure, notice, change
The role-routing adapter went through three policies. It started with difficulty tiers. Then I tried a strict policy: one fast premium model only, and only with a verified receipt. The logs showed that of 19 tasks under that policy, 1 completed with a verified receipt. The rest failed closed on timeouts or approval issues, and none was falsely marked compliant. The guard did its job. The policy was too brittle to use. I replaced it with role-based routing that follows my test-first playbook: a tester, an implementer and a read-only reviewer, each on its own tier, escalating only on high risk, a test dispute or a verified failed attempt.
Learn slowly, with my approval
A routing profile tracks the success rate per task signature with confidence bounds and waits for enough samples before it trusts them. Unobserved routes and environment failures are excluded. New policies go through shadow mode, then a canary, then full rollout, and full rollout needs my approval.
Result
The difficulty-routing adapter has the clearest number. Of its tasks where the model route was actually observed, ~81% ran on the free executor tier. Most tasks also skipped the planner.
The caveat, stated plainly: that share is slightly inflated, because verifier runs are logged on the free route by design. There is also no token or dollar baseline yet. The agent CLI build it drives exposes no per-call usage, so I cannot claim a cost saving, and I do not. The role-routing adapter records tokens but has no before-and-after comparison either.
What I’d do differently
I would set up a cost baseline before writing any routing logic, so the main claim could be measured in money instead of route shares. I would log verifier runs under their own route. I would also have tried the strict fast-only policy as a small canary first, since the receipts showed the friction within days.