Developer workflows

Run API and coding-agent stages in one reviewed pipeline

Requires: Gemini API + Claude Code + Codex + OpenAI API · This recipe needs every listed tool. Selecting one finds relevant workflows; it does not remove the other requirements.

Configure direct model calls alongside CLI jobs, inspect their request limits, and stop uncertain work before another dispatch.

Applies to
Requires: Gemini API + Claude Code + Codex + OpenAI API
Last verified
Reviewed by
Timothy Fehr

Use a direct model call for a stage that transforms supplied evidence into JSON. Keep coding-agent stages where their native behavior helps the task. The reference package can now supervise both in one run, with human decisions at scope, design, and acceptance.

This is optional. The default still uses synthetic answers. CLI account access and API credentials are configured separately; see the access reference. An IDE hosting a tool does not establish which account supplies that tool's access.

Give each stage a concrete job

The research recipe carries one retry contract through four model stages:

StageRecipientResult to inspect
Extract requirementsGemini APIClaims with source lines and unresolved questions
Propose a planClaude CodeA change tied to the selected requirements
Propose the configurationCodexA JSON candidate for the fixed checker
Review the evidenceOpenAI APIFindings and limitations after the checker ran

The role assignment is an experiment to evaluate. It does not establish a general ranking. The fixture's entire change is firstAttempt: 0 to firstAttempt: 1. Its checker evaluates JSON data and never executes generated code. The API stages have no browsing or repository tools.

Review the execution profile

Copy profiles/research.example.json from the download to a private local configuration. Prepare the CLI paths, account environment and separate workspace directory as described in CLI execution. The additional API entries look like this:

{
  "model": "REPLACE_WITH_AVAILABLE_RESPONSES_MODEL",
  "api_key_env": "PIPELINE_OPENAI_API_KEY",
  "max_output_tokens": 4096,
  "max_request_bytes": 65536
}

This is the value of apis["openai-api"]. Select an explicit model that supports the documented response format. Supply the key through your prepared process environment or secret manager. The profile contains its variable name; the coordinator reads the value only when dispatching that stage. It does not read saved chat or CLI credentials to obtain API access.

The reviewed profile shows the derived public endpoint, API family, requested model and bounds. Arbitrary endpoints, extra request parameters and undeclared API recipients are rejected. An API adapter omitted from apis waits for an explicit result import. Manual stages always wait for a person.

A sealed profile detects edits to its recorded configuration. It cannot prove which account a key belongs to, freeze a provider's model alias, or configure your network boundary. Check those against the approved account and data scope.

Start the mixed run

From the extracted package:

node cli.mjs new ./research-run --recipe research --execution native --profile ./research.local.json

Review the printed scope packet and record a named decision using its digest:

node cli.mjs decide ./research-run --gate scope --digest PRINTED_DIGEST --decision approve --by YOUR_NAME
node cli.mjs run ./research-run

The coordinator runs research and planning, then stops at design. Inspect the claims and plan before approving that gate with its newly printed digest. A subsequent run proposes the configuration, checks it and requests review. Acceptance requires another human decision over that evidence.

Keep the provider contracts visible

OpenAI Responses uses text.format for a JSON schema. Anthropic Messages uses output_config.format. This reference's Gemini adapter uses GenerateContent with generationConfig.responseFormat.text; Google labels that API family legacy. Interactions needs its own request builder and completion parser.

The clients send a schema containing the task's shape and allowed values. The complete schema stays in the prompt and local validator. Length, range and array-size constraints are checked locally after parsing; Anthropic documents limits on those constraints in the wire format. Successful HTTP transport alone cannot establish a valid task result.

Inspect failure before sending again

Each attempt records status, endpoint, model, request/response byte counts and hashes, plus safe usage counters when available. Provider error bodies are not printed or retained. Usage can be missing after a failure. The runtime limit, output-token allowance, byte caps and attempt budget do not impose an account spending cap.

Observed outcomeCoordinator behavior
Complete, valid resultContinue to the next eligible stage or human gate
HTTP 429, refusal, malformed JSON or invalid evidenceFail the stage; run alone does not retry it
RedirectRecord the status; do not follow it
Timeout, broken/oversized response, pending work, HTTP 408 or server errorRequire reconciliation and block new dispatch

For example, a connection can break after the provider accepted a request. The coordinator records remote_completion: "unknown". Inspect provider-side evidence and accounting before starting a newly approved run; this package cannot recover or prove the outcome automatically.

cancel or Ctrl-C aborts the local HTTP request and stops supervised CLI processes. Closing that connection does not prove remote computation stopped. Receipts and earlier attempts remain available for the human's review.

What goes wrong

A developer assumes a subscription funds direct API calls, supplies a model without the selected schema support, or copies an API key into a source packet. Another treats matching JSON as proof that the quoted requirement applies. Check the account, request contract and source evidence separately.

An invalid configuration may fail before any HTTP request exists. Correct it through a newly reviewed profile when its recorded values need to change. Do not remove the old run's receipts to regain its attempt budget.

How to check

node --test test.mjs executor.test.mjs api-executor.test.mjs

The HTTP tests use a local server and a synthetic key. They cover native request formats, mixed CLI/API handoffs, human pauses, independent reviews, redirects, failed responses, bounds, cancellation and validation. They consume no provider usage and do not establish credentialed integration.

For an account-backed trial, use the same small retry fixture. Keep the profile, returned model/usage, transport outcome and source/check evidence. Record that trial separately and compare the complete workflow's result and review effort.

Sources

  1. OpenAI: structured model outputs Tier 1 2026-09-08
  2. Anthropic: structured outputs Tier 1 2026-09-08
  3. Google: GenerateContent structured output Tier 1 2026-09-08
  4. Google: GenerateContent API reference Tier 1 2026-09-08