Developer workflows

Turn research into a checked code change

Requires: Gemini API + Claude Code + Codex + OpenAI API · This recipe needs every listed tool. Selecting one finds relevant workflows; it does not remove the other requirements.

Carry selected source evidence through planning, implementation, and review while keeping uncertainty and human decisions visible.

Applies to
Requires: Gemini API + Claude Code + Codex + OpenAI API
Last verified
Reviewed by
Timothy Fehr

A research summary becomes useful to implementation when every consequential claim points to source evidence and an acceptance condition. Carry the uncertainties forward so a later model cannot quietly turn them into facts.

This recipe demonstrates both coding-agent and direct-model stages. Direct API calls are separate products with separate access and accounting.

Start with selected evidence

The reference package includes a research recipe:

node cli.mjs new ./research-run --recipe research

The fixture's “research” source is a local retry contract. No internet search or live provider call occurs. This small corpus makes the chain of evidence easy to verify before using current external documentation.

The recorded role assignment is Gemini API for claim extraction, Claude Code for planning, Codex for a proposed change, and OpenAI API for review. Every role can be reconsidered through the recipe and its approved recipients.

Preserve a claim ledger

The research artifact contains claims with source path, line, and an exact supporting excerpt, plus unresolved questions. For example:

{
  "statement": "The first retry is numbered 1.",
  "source": {
    "path": "docs/retry-contract.md",
    "line": 2,
    "quote": "Attempt 1 waits 1000 ms."
  }
}

For external documentation, retain the official URL, access date, relevant version, and a permitted excerpt or source pointer. Distinguish documented behavior from an inference and from behavior observed in a test.

Do not treat a linked page as support for every nearby sentence. Read the specific supporting section and preserve contradictions rather than merging them into a confident recommendation.

Connect the design to a test

The planner receives the selected evidence and the original project snapshot. It should name the changed behavior and the check that will establish it.

In the fixture, the requirement “attempt 1 waits 1000 ms” maps to a deterministic case. The design gate shows that mapping before the writer proposes a value. Questions about jitter and server retry headers remain outside the fixture.

For a library migration, use the same pattern: pin the target version, identify the affected call sites, reproduce the old behavior, and specify the new acceptance check. A research report alone does not establish runtime behavior.

Choose the manual or automated research path

A developer can use ChatGPT research or Gemini with selected sources to prepare the same handoff. Change the research stage to manual and import the closed-schema JSON after inspecting the sources.

For programmatic extraction, use the documented model API. That does not recreate the corresponding chat research interface or a coding agent's tools. The reference supports explicit imports and optional supervised API requests. Its research profile configures Gemini API extraction, Claude Code planning, Codex implementation and OpenAI API review. Each stage keeps its own input and completion contract.

Prepare the CLI and API access separately, inspect the profile and limits, and approve scope before dispatch. The coordinator pauses again at design and acceptance. Local HTTP/subprocess tests verify these handoffs; they do not establish credentialed provider behavior.

Review both claims and behavior

The final reviewer receives the research ledger, plan, proposed configuration, and executed fixture checks. It can inspect the original source rather than relying only on the planner's summary.

At acceptance, check that important claims have support, the checks cover the changed behavior, and unresolved questions remain visible. A correct configuration can still be inappropriate if the source applies to a different version or product.

What goes wrong

An accurate quotation can still describe the wrong product version. A planner can turn an unresolved question into an assumed requirement. A passing unit check can miss the integration behavior that motivated the research. Review source applicability and runtime evidence together.

How to check

Run the package tests and change a quoted evidence string to a fabricated sentence. The research stage must fail before design or implementation starts.

For a real migration, compare the implementation against the selected version's primary documentation and run the relevant integration check in an isolated environment. Record what was observed, what is inferred, and what still needs a human decision.

Sources

  1. Google: structured outputs Tier 1 2026-09-08
  2. OpenAI: non-interactive Codex Tier 1 2026-09-08
  3. Anthropic: programmatic Claude Code Tier 1 2026-09-08