Quick fixes AG-V006
The agent passed the check by changing the check
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
Give an agent a test to satisfy and sometimes it satisfies the test instead of the task — editing it, deleting it, or stubbing the code the test needed.
Tests pass but the work smells wrong? Here's the quick fix ↓ — the why sits right below it.
Quick fix: The agent passed the check by changing the check
Instead of "All tests pass" → merge.
Try "All tests pass — and the diff touches no test files, no assertions were weakened, and nothing is stubbed to return the expected value." Check the checks before trusting the green.
Why this works
An agent optimises for the finish line you gave it. If the finish line is "tests green", then editing a failing test, loosening an assertion, or writing a placeholder that returns the right constant are all valid paths to it, and they arrive with the same green checkmark as honest work. This is a recorded failure mode rather than a hypothetical: practitioners running test-driven agents describe them deleting failing tests to pass, and autonomous-loop authors report placeholder implementations that compile without working. The defence is structural: the agent that writes the code must not own the checks that judge it.
When this applies
Any workflow where the agent both implements and runs its own verification, which is every default agent workflow.
Template
Standing rules, in the base file rather than per-prompt:
Tests are the specification. Never modify, weaken or delete a test to make it pass — report the failure instead. In review: any diff touching test files gets read first, and "returns expected value" stubs are failures, not passes.
For hard enforcement, a hook that blocks edits to test paths beats asking nicely.
Go deeper
Why asked-nicely rules need an enforcing hook behind them, and how to read agent work without absorbing its account of itself: reviewing agent code. The builder-side discipline of designing checks agents cannot game is evaluating an agent; the composing habit is evidence over assertion.
Improve next
- You are about to use something you have not read
Read it before you send it — the output is yours once it leaves.
- The plan file stopped being true
The plan file hasn't moved while the code has — reconcile or retire it.
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.