By context

A developer team's path through the guide

Shared methods · A shared method; linked tool guides explain the exact steps.

Choose product-specific onboarding, standardize evidence and access, and evaluate workflows against a measured baseline.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Start with the tools your developers actually use, then agree on common acceptance evidence and access boundaries. Keep each product's instructions and execution behavior explicit.

Choose a first complete task

Developers using a chat surface can start with ChatGPT, Gemini, or Claude projects, then hand off a clear task and source packet to the coding tool.

Standardize the evidence

A reviewed change should identify the requirement, source revision, diff, checks actually run, remaining findings, and a human acceptance decision. Use the same criteria across providers.

Keep stable repository guidance in the product's supported instruction mechanism. If the team maintains several files, nominate an owner for shared facts such as build commands and deployment boundaries. Verify consistency without assuming the files have identical loading rules.

Review permission configuration for every host and authentication path. An IDE-hosted agent may use the developer's own account. Check access and model references before comparing quota or cost.

Measure your own task classes

METR's early-2025 randomized study examined experienced open-source developers working in familiar repositories. It found a slowdown under those conditions, alongside a more positive perception of AI's effect. That result does not establish the effect of current models or unfamiliar tasks in your team.

Use it as a reason to measure. Compare recurring task classes on completion quality, cycle time, review time, defects, and usage. Record tool versions and access profiles. Include failures and abandoned attempts.

Avoid turning message volume or mandatory usage into a proxy for value.

Add pipelines after the handoff works

Start with an independent review. If it improves outcomes, try planning, implementation, and review.

Keep one owner of integration, separate candidate changes, and give independent reviewers the original evidence. A more elaborate recipe must justify its extra execution and human review work.

The reference package lets the team practice gates and recovery with synthetic data before configuring any live provider stage.

What goes wrong

A shared instruction file can drift from the build. A writer can weaken a test that then becomes its own evidence. Several agents can make incompatible changes to the same files. Extra review stages can add delay without finding useful defects. A provider switch can change who receives the source code.

How to check

Pick one recurring task and establish a baseline before the pilot. After several comparable runs, inspect accepted outcomes and the work needed to get there. Keep the approach for task classes where the evidence supports it, and revise or remove stages that add cost without useful improvement.

Use adoption measurement to keep usage, dependence, and outcomes separate.

Sources

  1. METR: early-2025 experienced developer productivity study Tier 2 2026-09-08