Developer workflows

Recover a pipeline without losing its boundaries

Shared methods · A shared method; linked tool guides explain the exact steps.

Handle failed output, denied decisions, cancellation, retries, and limited budgets while preserving the run's evidence.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

A resumable pipeline needs a record of what actually completed and what remains uncertain. A process restart does not prove success, and it does not create a new permission to repeat an action.

The reference package demonstrates recovery for local artifacts. It deliberately has no external side-effect executor.

Distinguish the failure before retrying

OutcomeNext action
Invalid JSON, schema, or evidenceMark the stage failed; inspect the cause before a bounded retry.
Process exit failure or missing final resultKeep the output as unsuccessful; do not open the next gate.
Permission deniedNarrow the task or obtain the concrete decision required by the coordinator.
Human denialPreserve the denial across restarts.
Check failureReturn the evidence to acceptance; request a bounded revision or stop.
Interrupted running stageReconcile its outcome before deciding whether another run is appropriate.
Unknown result of a remote actionInspect the target system before considering any repeat.

A timeout after creating a remote resource is different from a timeout before dispatch. A local event log cannot promise exactly-once external effects.

Use the recorded state

node cli.mjs status ./my-run
node cli.mjs retry ./my-run --stage correctness
node cli.mjs run ./my-run

retry only resets a failed stage that has attempts remaining. It does not override a denied human gate or a stage marked needs_reconciliation.

Artifacts are immutable and hash-checked on load. Recipe and baseline changes invalidate the run. Imported results must carry the current request digest, which includes the stage attempt. A previous attempt's answer cannot be silently attached to a new request.

The runner serializes mutations with a lock. After a crash that leaves the lock behind, inspect its recorded PID and confirm that the process has stopped. Inspect the state and reconcile interrupted work before removing the lock. Never unlock a run merely because it has been quiet for a while.

Reserve shared limits before dispatch

Each recipe declares concurrency, per-stage and total attempts, output bounds, a run deadline, and a revision limit. The runner reserves the maximum output allowance for each stage before dispatching an independent batch.

The example charges each attempt against its reserved allowance. A failed attempt does not restore that allowance. This conservative policy is easy to inspect and prevents parallel work from silently exceeding the total.

The deadline includes time spent waiting for a human. If it expires, create a newly approved run. Adjust the recipe before starting when a review is expected to take longer.

These are coordinator dispatch and artifact limits. Import mode cannot stop a native process launched separately or enforce its monetary spend. Configure and verify deadline, process, tool, and account controls in that execution environment. Report incomplete usage as incomplete.

Cancel with a reconciliation list

node cli.mjs cancel ./my-run --by YOUR_NAME

Cancellation stops future coordinator work and rejects later imports. It also records stages that were waiting for externally produced output.

Stop any native sessions you launched separately and inspect whether they created files or external effects. The reference runner does not own those processes and cannot terminate them. Do not claim that cancelling its record cancelled a provider job.

For an automated executor, cancellation should terminate the process tree where supported, preserve the termination result, and report any work that could not be stopped. Add those checks when adding that executor.

Declare fallback recipients in advance

A failed provider must not cause silent routing to a different organization. The reference validates each adapter against the recipe's recipient list.

Even an allowed alternative may use a different model, billing route, context strategy, or retention policy. Create a new reviewed recipe when those material conditions change. Preserve the failed run as evidence of why the change was needed.

What goes wrong

Retrying every error can repeat a remote action, consume the whole allowance, or repeatedly encounter the same permission denial. Restoring a run from an old file can also discard a human rejection. Preserve the recorded decisions and reconcile uncertain outcomes.

How to check

Run node --test test.mjs. The suite exercises denial persistence, stale approvals, exhausted output allowances, expired deadlines, cancelled imports, and interrupted running states. None of these cases should advance automatically into acceptance.

Then practice recovery on the synthetic fixture: deny scope, fail an import, retry once, and cancel another run. The status report should explain each outcome without relying on the operator's chat history.

Sources

  1. OpenAI: non-interactive Codex Tier 1 2026-09-08
  2. Anthropic: programmatic Claude Code Tier 1 2026-09-08
  3. Google: Gemini CLI headless mode Tier 1 2026-09-08