Run CLI stages between human decisions
Requires: Claude Code + Codex + Gemini CLI · This recipe needs every listed tool. Selecting one finds relevant workflows; it does not remove the other requirements.
This page covers tools outside your selection. You can still read it. Find matching guides
Configure Claude Code, Codex, and Gemini CLI, review their execution profile, and let the local coordinator launch and supervise eligible stages.
Once a handoff has a clear input, output contract, and acceptance check, a coordinator can launch the next eligible coding tool. Human decisions still belong at scope, design, and acceptance.
The downloadable reference now includes optional native CLI execution. Its concrete task stays small: correct firstAttempt in a retry configuration, check the expected delays, and review the evidence. The website runs none of these commands.
Choose an execution mode
| Mode | What happens |
|---|---|
fixtures, the default | Recorded synthetic outputs exercise the complete recipe without provider calls. |
import | A person launches their tool and supplies output plus observed transport status. |
native | Configured CLI stages run as supervised child processes. Optional API configuration enables direct requests; manual and unconfigured API stages wait for imports. |
Codex and Claude Code have native programmatic interfaces. Gemini CLI's JSON envelope contains an inner answer that still needs parsing and validation. A model API call does not provide a coding agent's tools or configuration.
Prepare the execution profile
For the mixed research recipe, see API execution. The CLI-only implementation profile below needs no API configuration.
Extract the package and copy profiles/native.example.json to a private local configuration file. Replace all placeholders. Select an installed executable, available model, declared tool version, and the environment variable names each tool needs for its account.
On Windows, choose an actual .exe. For an npm CLI, choose node.exe plus the installed JavaScript entrypoint. Shell shims such as codex.ps1 and gemini.cmd are rejected. The example paths illustrate the fields; inspect your installation.
Create a separate workspace_root directory for native attempts. Describe the prepared environment in isolation_note: which files, credentials and network access the processes actually receive. That note is documentation, not a sandbox configuration. Set up and verify the boundary before approving a run.
Gemini CLI needs an explicit reviewed policy file. The package includes a deny-tools example for this JSON-only task. Confirm that it is effective in your installed version: standard admin policies can cause supplemental --admin-policy files to be ignored. A policy file does not disable startup extensions or provide operating-system isolation.
Run the checked configuration task
Use the existing implementation recipe:
node cli.mjs new ./native-run --recipe implementation --execution native --profile /absolute/path/to/my-profile.json
This prints the scope packet before any model stage runs. Review the source files, recipients, limits, requested models, inherited variable names, and sealed launcher/entrypoint/policy hashes. Credential values are not part of that packet.
Record your decision against its printed digest, then advance:
node cli.mjs decide ./native-run --gate scope --digest PRINTED_DIGEST --decision approve --by YOUR_NAME
node cli.mjs run ./native-run
Claude Code proposes a plan and the run stops for the design decision. Inspect the plan and transport receipt. After approving the new design digest, run again: Codex proposes the configuration, the fixed checker verifies it, and Claude Code and Gemini CLI independently review the same evidence. Acceptance requires another human decision.
The candidate changes firstAttempt from 0 to 1. Expected delays for attempts 1, 2 and 6 are 1000, 2000 and 30000 ms; attempts 0 and 1.5 are invalid. The checker evaluates JSON data. This recipe does not apply repository edits.
Inspect failures and cancellation
The runner records observed exit status, requested model, elapsed time, and output byte counts and hashes. Actual model identity can remain unknown. Successful transport must still pass envelope, schema, and evidence checks. A failed stage or reconciliation requirement makes run exit with code 1.
Attempts have a runtime limit, stdout/stderr caps, a shared byte reservation, and the run's overall deadline. These limits do not impose a monetary cap on your account.
Cancel from another terminal:
node cli.mjs cancel ./native-run --by YOUR_NAME
The active coordinator observes the request and stops its owned process tree. Ctrl-C also requests shutdown. Local process closure does not prove that a remote request or escaped service stopped. Interrupted or unconfirmed work needs reconciliation; the runner does not automatically retry it or kill a PID copied from an old run.
What goes wrong
Changing a sealed profile invalidates approval. Updating a launcher or policy stops dispatch until you create a newly approved run. Launcher hashes do not pin every dependency, account setting or remote service. An IDE host does not establish whose subscription is used; check the tool's actual account.
Treating a fresh directory as filesystem isolation exposes more than the approved packet. Treating process exit as acceptance skips the task checks. The package README explains these boundaries and how to inspect failures.
How to check
From the extracted package, run:
node --test test.mjs executor.test.mjs api-executor.test.mjs
These tests use real local child processes with synthetic replies. They check human pauses, concurrent independent reviews, environment filtering, changed launchers, failed exit codes, oversized output, timeout, and cancellation from a second CLI process.
This pass ran on Windows with Node.js 24.13.0. Codex CLI 0.153.4 and Claude Code 2.1.263 flags were checked locally; Gemini CLI was unavailable. Credentialed provider behavior and POSIX termination were not exercised here. Before using your account, review the profile and run the same synthetic task with your installed tools. Keep that evidence separate from the package's offline tests.
Sources
- OpenAI: non-interactive Codex Tier 1 2026-09-08
- Anthropic: programmatic Claude Code Tier 1 2026-09-08
- Google: Gemini CLI policy engine Tier 1 2026-09-08
- Node.js: child processes Tier 1 2026-09-08
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.