plan → execute

Design

Six roles

The original version of this setup called it five. It listed six.

Role Where it lives Model Why it exists
Planner skills/make-plan opus (fable when the design is genuinely open) Synthesis: architecture, phase design, risk, resolving competing approaches
Planning researcher agents/planning-researcher.md sonnet Breadth-first codebase search, so the planner is not paying planner rates to grep
Orchestrator skills/execute-plan sonnet Sequencing, delegating, reviewing, filtering findings
Implementer agents/implementer.md sonnet Writes the code and the tests. Most tokens go here
Quick implementer agents/quick-implementer.md haiku Mechanical, exactly-specified work. Refuses anything else
Reviewer agents/plan-reviewer.md opus Read-only gate at risk points and at the end

“Most tokens go here” is measured, not assumed: on a real /execute-plan run (2.1.238), 79% of output tokens came out of the workers, on their pinned models — scripts/verify-models.py prints the per-tier split from the transcripts, deduped and grouped by each worker’s recorded agentType.

Why subagents and not skills for the workers

Skills can set model: for the turn — literally for the turn, which is its own trap, see Gotchas — and can even fork into a subagent with context: fork. But workers need isolated context as much as they need a model pin: verbose test output, file reads, and failed attempts stay inside the worker, and only a short report crosses back. Subagents give both. That isolation is the entire reason a long implementation phase does not swamp the orchestrator’s context.

The corollary is that “report concisely” in the worker prompts is a live constraint, not boilerplate. If workers write long reports the isolation buys nothing.

Why the plan is a file

A plan held in a conversation dies with that conversation, and survives compaction only by luck. On disk it can be executed in a fresh session that never loads the planning transcript — cheaper, and cleaner, because the executor sees requirements instead of the deliberation that produced them.

That only works if the plan is genuinely self-contained, which is what the plan-file contract enforces: base commit, goal, success criteria, non-goals, constraints, global verification commands, and per phase a risk level, verified evidence, required behaviour, and “done when”.

Claude Code’s own plan mode also persists to disk now. The plansDirectory setting — “Custom directory for plan files, relative to project root. If not set, defaults to ~/.claude/plans/” — points it at the repository. Set it to plans and Shift+Tab plan mode writes to the same place /make-plan does. /make-plan still earns its keep by controlling the filename, capping research delegation, and filling the contract.

They are alternatives, though, not layers. Plan mode injects its own workflow, and that injection outranks a skill body: research gets forced onto Explore, which inherits the session model, and writes are fenced to the harness’s own plan path. Run /make-plan from plan mode and you get a plan outside the repository produced by an expensive model doing its own grepping — the two failures this whole design exists to prevent, at once. So /make-plan checks and refuses. See Gotchas.

The phase packet

The orchestrator does not hand a worker the whole plan, and does not hand it the phase alone.

Whole plan: the worker re-reads phases 1–4 to implement phase 5, every time.

Phase alone: the worker satisfies the phase and breaks the product, because it cannot see what the phase is for. This is the flaw in the obvious token-saving advice.

The packet is the middle: overall goal, the success criteria this phase serves, the phase verbatim, the constraints and non-goals that bind it, dependencies and deviations from earlier phases, files renamed and APIs changed so far, and the pre-existing dirty files it must not touch. The orchestrator decides what is relevant — that is what an orchestrator is for.

The git discipline

This closes a real hole rather than optimising anything.

git diff compares the working tree against the index, so it shows nothing for staged changes and nothing for untracked files. An orchestrator told to “read the actual diff” can therefore approve a phase without ever seeing the new migration, the new component, or the new test file it added.

So: record git status --short before the phase, record it again after, and use the difference to find every file the phase touched. Review tracked changes with git diff HEAD -- <files>. Read every new untracked file directly. Ignore the orchestrator’s own plan-file status edits. The reviewer gets the same list, including the untracked files, because it has the same blind spot.

The final integration gate

Per-phase review is not enough: a phase can pass in isolation while the combination fails. After the last phase, the orchestrator runs the plan’s global verification commands against the accumulated tree, then one reviewer pass over the whole change set — including every newly created file — against the goal, the success criteria, and the non-goals. Findings route back, affected checks re-run, and only then is it done.

Then one more thing, cheap because the material is already collected: if the run recorded deviations or consciously declined findings, the orchestrator closes with a one-line-each retro on the ones that carry durable repository knowledge — a plan assumption the codebase contradicted, a missing test wrapper, a convention no document states — and names where each belongs: CLAUDE.md, a bin/ wrapper, the plan template. Suggestions only; the orchestrator cannot create files, and the user decides what gets recorded. The deviation log preserves what happened; the retro is for the subset that should change what the next session reads before it starts.

Coverage lives in the reviewer, filtering lives in the orchestrator

The reviewer reports every plausible finding in its categories, tagged with a severity and a confidence, and filters nothing. The orchestrator decides what matters, routes those back, and records in the plan file anything it consciously declines.

This split exists because a severity filter inside the reviewer reads as an instruction to withhold, and current models follow it literally: they investigate just as hard, find the bug, and decline to mention it. Precision goes up, recall goes down, and real defects vanish silently. See Prompting.

The reviewer is also told not to write the fix — file, symbol, evidence, impact, smallest conceptual correction. Two reasons: an expensive model producing patch text is expensive output for work the cheap model is about to do anyway, and a pasted patch collapses the role split that makes the review independent.

What the orchestrator is actually prevented from doing

“The orchestrator writes no code” is two different guarantees, and only one of them is enforced.

Write is withheld through the skill’s disallowed-tools, so it cannot create a file at all — no new component, no new test, no new migration appearing from the orchestrator tier. That part is the harness’s job and it holds.

Edit it keeps, because the plan file needs editing for phase status and deviations, and frontmatter cannot scope a tool to one path. So “the plan file only” is a rule the orchestrator holds itself to. If you need that enforced rather than instructed, a PreToolUse hook on Edit is the instrument — with the caveat that hooks are dropped for plugin-packaged skills, so it protects only the install.sh path.

An allowlist would be the stronger-looking option and is the wrong one here: omitting ToolSearch makes the deferred SendMessage unloadable, which silently breaks resuming the same implementer inside a phase.

Delegate, then commit to the delegation

An orchestrator that re-derives what a subagent already reported pays for the same work twice at the more expensive tier. The rule is: read the diff to verify the work, not to repeat it. Corrections inside a phase resume the same implementer (it keeps its context); the next phase gets a fresh one; reviewers are always fresh.

When not to use this

A fresh subagent starts cold: no conversation, no prior file reads, no warm cache. So orchestrator → implementer → orchestrator reads the diff can pay twice to read the same problem. For a three-file feature, one continuous Sonnet session is often both cheaper and better oriented.

Isolation earns its cost when the implementation produces a lot of output, the run is long, or the phases are genuinely independent.

Work Workflow
Small, deterministic edit Sonnet directly; batched Haiku for mechanical sets
Normal feature, clear multi-file bug Sonnet directly, optionally one closing review
Long-running or context-heavy feature /make-plan → /execute-plan
Migration, auth, security, public API, AI integration Opus plan → /execute-plan → review gates
Genuinely open-ended architecture Fable plan → same execution path
Plan turns out to be wrong Stop, fix the plan with the big model, record the deviation, resume

A typo, an obvious one-file bug, or a fully specified UI tweak does not need a plan file at all.

Stopping a stuck implementer

The tempting control is maxTurns, and it is the wrong one twice over. Even if it worked it would be wrong at small values: a turn is every agentic step — read the packet, grep, read the service, read the test, edit, run, read the failure, fix — so a low cap truncates correct work mid-phase and leaves half-written code behind. And it does not work; see below.

The failure actually worth stopping is edits that gain no information. So the implementer is told: if two materially different repair attempts produce the same failure and nothing was learned between them, stop and report the failing command, the hypotheses tested, what changed between attempts, the evidence, and the most likely next investigation. Three red tests in a row is often just a correct sequence — the reproduction test failing as designed, then an edge case surfacing, then a stale fixture.

Do not reach for maxTurns as the backstop either. Measured on a subagent that declared maxTurns: 15: it ran 65 turns. The key is accepted and ignored, so the instruction above is the whole mechanism — there is no hard limit behind it.

Why the read-only agents carry no turn limit

planning-researcher and plan-reviewer had maxTurns: 15 and it was removed. The reasoning above applies to them with more force, not less.

The first reason is that it was not doing anything. A researcher delegated five bugs spanning two repositories was reported back to the planner as Done · 21 tool uses · 49.8k tokens · 1m 16s — and its subagent transcript shows it was at turn 33 of an eventual 65 when that happened. A cap of 15 did not stop it at 15, or at all. Inert config is worse than no config: it reads like a control, so you stop looking for the real one.

What the planner received as the agent’s “final answer” was its interleaved narration — “Found AppSetting, EnforceAppSettings middleware… let’s dig in” — because the run was cut mid-investigation, before the turn that writes the report. The planner said the answer kept truncating and went back to grepping the codebase itself at planner rates. Its follow-up then resumed the agent in the background, where it ran another 32 turns that nothing was waiting for.

So the shape of the failure is real even though the cap was not its cause: an agent that searches first and reports last loses everything if it is stopped early, and the visible symptom is not “stopped early” but “the agent lost its answer”. The defence is to size the delegation so it finishes — see the research section of make-plan — not to add a second stopping condition.

A cap is also the wrong place to economise. Research is already the cheap tier — planning-researcher runs Sonnet at medium precisely so breadth-first grepping does not run at planner rates. That is where the saving comes from. A turn cap on top of it buys nothing and buys it at the price of the report.

The risk a cap guards against is a runaway loop, and what makes a loop expensive is repeated writes. Neither agent has Edit or Write. Both keep Bash, because breadth-first grep and git log are the job — a measured research run made thirty-odd Bash calls and a handful of Reads. Bash is not a sandbox, and permissionMode: plan gates it only in a local install; a plugin install drops that key (see gotchas.md).

So the guarantee here is narrow, and worth stating exactly rather than rounding up to “read-only”: these agents will not sit in a write-test-write loop, so a search that runs ten turns too long costs ten cheap turns and nothing else. A truncated report costs the whole delegation.

Cap agents that edit. Let agents that only look, finish. And before trusting any cap, check a subagent transcript under ~/.claude/projects/<project>/<session>/subagents/ and count the turns it actually took.

Effort

effort is available on both skills and agents: low, medium, high, xhigh, max, or an integer. The top levels are model-gated.

The split that matters is orchestrator versus implementer. Finding the next phase, assembling a packet, and checking a box do not need the same reasoning budget as writing the code. So the orchestrator skill runs medium and implementer overrides back up to high. Without the explicit effort: high on the implementer it would inherit the orchestrator’s medium.

The orchestrator’s medium is conditional, and the condition is invisible while it holds. Skill frontmatter applies for the current turn only, so medium survives a run only while every delegation stays in the foreground: a backgrounded worker’s completion notification starts a new turn at the session’s effort. Measured medium→xhigh at the first notification in three runs (2.1.239, 2.1.251, 2.1.259). Two defences, both in the settings snippet — CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 in the project env keeps delegations in the foreground, and the session should be set to Sonnet at medium before /execute-plan is invoked (modelSettings.claude-sonnet-5.effortLevel makes that standing). Both are overridden by CLAUDE_CODE_EFFORT_LEVEL if it is set. See Gotchas.

Agent effort is fixed per definition — there is no per-invocation knob. So “bump the final review to xhigh” is not something the orchestrator can do; it would need a second agent definition. The reviewer stays at high, and the final gate is deepened by telling it in the delegation that this is the whole accumulated change set against every success criterion.