Business and governance

Measure outcomes alongside AI adoption

Shared methods · A shared method; linked tool guides explain the exact steps.

Distinguish usage, reliance, and value, and compare complete developer tasks instead of treating message volume as productivity.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Track three different questions: who uses the tool, which work depends on it, and whether that work has improved. A usage dashboard answers the first question more readily than the other two.

This method applies across Claude, ChatGPT, Gemini, their coding counterparts, and pipelines that combine them.

Define the outcome before counting activity

For a recurring bug fix, record time to a reviewed change, reviewer effort, missed regressions, and the checks required for acceptance. For documentation, record verified claims, corrections, and time to an approved article.

Keep a denominator and comparable task class. “Review time fell” needs the number and kind of changes reviewed. More completed tickets can reflect smaller tickets rather than a better development workflow.

Keep usage and reliance visible

Account activation shows provisioning. Activity shows that a tool was used. Reliance asks what work now routes through it and what would happen if it were unavailable.

These are useful operational measures. They become misleading when presented alone as proof of better outcomes.

For a pipeline, count stages, retries, manual handoffs, and the human time spent resolving findings. A short final answer can hide a large amount of failed work.

Treat self-report as one source of evidence

METR's early-2025 randomized study found that experienced developers on familiar open-source repositories could feel faster while taking longer with the available AI tools. Its task population and tool generation limit how broadly the result can be applied.

The practical lesson is to pair satisfaction with observed outcomes. Satisfaction can explain willingness to adopt a tool; it does not establish the effect on quality or elapsed work.

Try a small comparison

Choose a recurring task class and agree on acceptance checks. Compare the existing method with a single-tool approach and, when useful, a cross-provider pipeline. Keep the source revision and task criteria comparable.

Record actual model and tool versions where observable. Mark unknown metadata and incomplete usage explicitly. Include review, rework, failures, and waiting time, rather than only model response time.

Review the results by task class. A tool can help unfamiliar-code exploration while adding work on another task; that is something to measure, not assume.

What goes wrong

Message targets encourage more messages. A pilot can select only easy wins. Before-and-after comparisons can confound changed team practices with the tool's effect. A satisfaction score can obscure additional review load. Averages can hide a few costly failed runs.

Use comparable samples and describe the remaining uncertainty.

How to check

Take an existing dashboard and label every measure as usage, reliance, or outcome. Add at least one outcome measure for the main task the rollout is intended to improve.

Inspect several accepted and failed tasks with their reviewers. The result should explain where the tool helps, where it adds work, and what evidence would change the team's decision.

Sources

  1. METR: early-2025 experienced developer productivity study Tier 2 2026-09-08