Run API and coding-agent stages in one reviewed pipeline
Requires: Gemini API + Claude Code + Codex + OpenAI API · This recipe needs every listed tool. Selecting one finds relevant workflows; it does not remove the other requirements.
This page covers tools outside your selection. You can still read it. Find matching guides
Configure direct model calls alongside CLI jobs, inspect their request limits, and stop uncertain work before another dispatch.
Use a direct model call for a stage that transforms supplied evidence into JSON. Keep coding-agent stages where their native behavior helps the task. The reference package can now supervise both in one run, with human decisions at scope, design, and acceptance.
This is optional. The default still uses synthetic answers. CLI account access and API credentials are configured separately; see the access reference. An IDE hosting a tool does not establish which account supplies that tool's access.
Give each stage a concrete job
The research recipe carries one retry contract through four model stages:
| Stage | Recipient | Result to inspect |
|---|---|---|
| Extract requirements | Gemini API | Claims with source lines and unresolved questions |
| Propose a plan | Claude Code | A change tied to the selected requirements |
| Propose the configuration | Codex | A JSON candidate for the fixed checker |
| Review the evidence | OpenAI API | Findings and limitations after the checker ran |
The role assignment is an experiment to evaluate. It does not establish a general ranking. The fixture's entire change is firstAttempt: 0 to firstAttempt: 1. Its checker evaluates JSON data and never executes generated code. The API stages have no browsing or repository tools.
Review the execution profile
Copy profiles/research.example.json from the download to a private local configuration. Prepare the CLI paths, account environment and separate workspace directory as described in CLI execution. The additional API entries look like this:
{
"model": "REPLACE_WITH_AVAILABLE_RESPONSES_MODEL",
"api_key_env": "PIPELINE_OPENAI_API_KEY",
"max_output_tokens": 4096,
"max_request_bytes": 65536
}
This is the value of apis["openai-api"]. Select an explicit model that supports the documented response format. Supply the key through your prepared process environment or secret manager. The profile contains its variable name; the coordinator reads the value only when dispatching that stage. It does not read saved chat or CLI credentials to obtain API access.
The reviewed profile shows the derived public endpoint, API family, requested model and bounds. Arbitrary endpoints, extra request parameters and undeclared API recipients are rejected. An API adapter omitted from apis waits for an explicit result import. Manual stages always wait for a person.
A sealed profile detects edits to its recorded configuration. It cannot prove which account a key belongs to, freeze a provider's model alias, or configure your network boundary. Check those against the approved account and data scope.
Start the mixed run
From the extracted package:
node cli.mjs new ./research-run --recipe research --execution native --profile ./research.local.json
Review the printed scope packet and record a named decision using its digest:
node cli.mjs decide ./research-run --gate scope --digest PRINTED_DIGEST --decision approve --by YOUR_NAME
node cli.mjs run ./research-run
The coordinator runs research and planning, then stops at design. Inspect the claims and plan before approving that gate with its newly printed digest. A subsequent run proposes the configuration, checks it and requests review. Acceptance requires another human decision over that evidence.
Keep the provider contracts visible
OpenAI Responses uses text.format for a JSON schema. Anthropic Messages uses output_config.format. This reference's Gemini adapter uses GenerateContent with generationConfig.responseFormat.text; Google labels that API family legacy. Interactions needs its own request builder and completion parser.
The clients send a schema containing the task's shape and allowed values. The complete schema stays in the prompt and local validator. Length, range and array-size constraints are checked locally after parsing; Anthropic documents limits on those constraints in the wire format. Successful HTTP transport alone cannot establish a valid task result.
Inspect failure before sending again
Each attempt records status, endpoint, model, request/response byte counts and hashes, plus safe usage counters when available. Provider error bodies are not printed or retained. Usage can be missing after a failure. The runtime limit, output-token allowance, byte caps and attempt budget do not impose an account spending cap.
| Observed outcome | Coordinator behavior |
|---|---|
| Complete, valid result | Continue to the next eligible stage or human gate |
| HTTP 429, refusal, malformed JSON or invalid evidence | Fail the stage; run alone does not retry it |
| Redirect | Record the status; do not follow it |
| Timeout, broken/oversized response, pending work, HTTP 408 or server error | Require reconciliation and block new dispatch |
For example, a connection can break after the provider accepted a request. The coordinator records remote_completion: "unknown". Inspect provider-side evidence and accounting before starting a newly approved run; this package cannot recover or prove the outcome automatically.
cancel or Ctrl-C aborts the local HTTP request and stops supervised CLI processes. Closing that connection does not prove remote computation stopped. Receipts and earlier attempts remain available for the human's review.
What goes wrong
A developer assumes a subscription funds direct API calls, supplies a model without the selected schema support, or copies an API key into a source packet. Another treats matching JSON as proof that the quoted requirement applies. Check the account, request contract and source evidence separately.
An invalid configuration may fail before any HTTP request exists. Correct it through a newly reviewed profile when its recorded values need to change. Do not remove the old run's receipts to regain its attempt budget.
How to check
node --test test.mjs executor.test.mjs api-executor.test.mjs
The HTTP tests use a local server and a synthetic key. They cover native request formats, mixed CLI/API handoffs, human pauses, independent reviews, redirects, failed responses, bounds, cancellation and validation. They consume no provider usage and do not establish credentialed integration.
For an account-backed trial, use the same small retry fixture. Keep the profile, returned model/usage, transport outcome and source/check evidence. Record that trial separately and compare the complete workflow's result and review effort.
Sources
- OpenAI: structured model outputs Tier 1 2026-09-08
- Anthropic: structured outputs Tier 1 2026-09-08
- Google: GenerateContent structured output Tier 1 2026-09-08
- Google: GenerateContent API reference Tier 1 2026-09-08
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.