Turn a working handoff into a developer pipeline
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
Automate a known stage, keep its evidence, and retain human decisions at the points where they matter.
A useful workflow has a result you can recognize. A pipeline automates some of the steps that produce it. Begin with a task you can already complete by hand, so you can tell whether automation improves the result.
Use a small code change as the exercise: correct a retry delay, review the diff, and decide whether to accept it. The same approach applies to a technical research brief or a report built from CI data.
Do the handoff manually first
Give the implementer a task, relevant files, and acceptance checks. After the change, collect the exact diff, test output, source revision, and unresolved questions. Give a reviewer the original evidence and ask for specific findings.
You can do the review in Claude, ChatGPT, or Gemini by supplying the packet. A coding agent can inspect a checkout when its access permits. Choose the available interface and record where execution actually happens.
An example handoff is:
Task: First retry must wait 1 second; retain the 30-second cap.
Inputs: retry.js and caller.js at the recorded base revision.
Change: Adjust the exponent while keeping attempts numbered from one.
Checks: Original fails attempt 1; candidate passes attempts 1, 2, and 6.
Open decision: Behavior for zero or negative attempts.
Deliverable: Findings with file, line, failure scenario, and suggested check.
Keep source files with this summary. A reviewer needs a way to test its claims.
Automate one reliable stage
Start with deterministic collection: export the diff, copy the approved input files, run known checks, and create an artifact manifest. Keep the review manual for this first version. That alone removes repeated transfers.
Then replace the manual review stage with a native adapter if it is useful. For example, Codex exposes a non-interactive command with structured output. Parse the tool's envelope, validate the answer, and return the report to the same human review point. The programmatic guides describe each product's interface and prerequisites.
The manual and automated versions must meet the same acceptance criteria. Automation does not lower the standard because a stage is inconvenient to check.
Add another provider deliberately
Try a second reviewer against the same frozen diff. Keep its first report independent of the other review so it can look for a different failure. Combine findings by their evidence and observable consequence.
For each added stage, measure whether it finds useful problems and how much review work it creates. Several agreeing models can share the same mistaken assumption. Test a disagreement or bring the actual decision to the human.
Keep the coordinator in charge
The coordinator owns the stages, allowed recipients, limits, and decision records. A model can propose a new step; it cannot grant itself the authority to execute it. Native tool approvals and workflow acceptance answer different questions and need the appropriate enforcement.
Read artifacts and human decisions before connecting implementation or publication steps. A person can approve the concrete next action once its result and target are visible.
What goes wrong
A script forwards an entire conversation because no handoff format was defined. A reviewer checks the author's summary instead of the changed code. A retry silently starts work against a newer checkout. A missing provider is replaced with another recipient the user never selected.
Give each stage explicit inputs, outputs, limits, and failure behavior. Start with the read-only review recipe before adding writers.
How to check
Run one task manually and through the automated version. Compare the actual inputs, findings, checks, and human decisions. Introduce a known defect and confirm it reaches the final review packet.
Cancel a stage and deny an approval. Neither action should be interpreted as success. Record elapsed time, manual corrections, and usage before deciding whether the extra automation earns its place.
Sources
- OpenAI: non-interactive Codex Tier 1 2026-09-08
- Anthropic: approvals and user input Tier 1 2026-09-08
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.