Design
Six roles
The original version of this setup called it five. It listed six.
| Role | Where it lives | Model | Why it exists |
|---|---|---|---|
| Planner | skills/make-plan |
opus (fable when the design is genuinely open) | Synthesis: architecture, phase design, risk, resolving competing approaches |
| Planning researcher | agents/planning-researcher.md |
sonnet | Breadth-first codebase search, so the planner is not paying planner rates to grep |
| Orchestrator | skills/execute-plan |
sonnet | Sequencing, delegating, reviewing, filtering findings |
| Implementer | agents/implementer.md |
sonnet | Writes the code and the tests. Most tokens go here |
| Quick implementer | agents/quick-implementer.md |
haiku | Mechanical, exactly-specified work. Refuses anything else |
| Reviewer | agents/plan-reviewer.md |
opus | Read-only gate at risk points and at the end |
“Most tokens go here” is measured, not assumed: on a real /execute-plan run
(2.1.238), 79% of output tokens came out of the workers, on their pinned
models — scripts/verify-models.py prints the per-tier split from the
transcripts, deduped and grouped by each worker’s recorded agentType.
Why subagents and not skills for the workers
Skills can set model: for the turn — literally for the turn, which is its own
trap, see Gotchas — and can even fork into a subagent with
context: fork. But workers need isolated context as much as they need a
model pin: verbose test output, file reads, and failed attempts stay inside the
worker, and only a short report crosses back. Subagents give both. That isolation
is the entire reason a long implementation phase does not swamp the
orchestrator’s context.
The corollary is that “report concisely” in the worker prompts is a live constraint, not boilerplate. If workers write long reports the isolation buys nothing.
Why the plan is a file
A plan held in a conversation dies with that conversation, and survives compaction only by luck. On disk it can be executed in a fresh session that never loads the planning transcript — cheaper, and cleaner, because the executor sees requirements instead of the deliberation that produced them.
That only works if the plan is genuinely self-contained, which is what the plan-file contract enforces: base commit, goal, success criteria, non-goals, constraints, global verification commands, and per phase a risk level, verified evidence, required behaviour, and “done when”.
Claude Code’s own plan mode also persists to disk now. The plansDirectory
setting — “Custom directory for plan files, relative to project root. If not
set, defaults to ~/.claude/plans/” — points it at the repository. Set it to
plans and Shift+Tab plan mode writes to the same place /make-plan does.
/make-plan still earns its keep by controlling the filename, capping research
delegation, and filling the contract.
They are alternatives, though, not layers. Plan mode injects its own workflow,
and that injection outranks a skill body: research gets forced onto Explore,
which inherits the session model, and writes are fenced to the harness’s own plan
path. Run /make-plan from plan mode and you get a plan outside the repository
produced by an expensive model doing its own grepping — the two failures this
whole design exists to prevent, at once. So /make-plan checks and refuses. See
Gotchas.
The phase packet
The orchestrator does not hand a worker the whole plan, and does not hand it the phase alone.
Whole plan: the worker re-reads phases 1–4 to implement phase 5, every time.
Phase alone: the worker satisfies the phase and breaks the product, because it cannot see what the phase is for. This is the flaw in the obvious token-saving advice.
The packet is the middle: overall goal, the success criteria this phase serves, the phase verbatim, the constraints and non-goals that bind it, dependencies and deviations from earlier phases, files renamed and APIs changed so far, and the pre-existing dirty files it must not touch. The orchestrator decides what is relevant — that is what an orchestrator is for.
The git discipline
This closes a real hole rather than optimising anything.
git diff compares the working tree against the index, so it shows nothing
for staged changes and nothing for untracked files. An orchestrator told to
“read the actual diff” can therefore approve a phase without ever seeing the new
migration, the new component, or the new test file it added.
So: record git status --short before the phase, record it again after, and use
the difference to find every file the phase touched. Review tracked changes with
git diff HEAD -- <files>. Read every new untracked file directly. Ignore the
orchestrator’s own plan-file status edits. The reviewer gets the same list,
including the untracked files, because it has the same blind spot.
The final integration gate
Per-phase review is not enough: a phase can pass in isolation while the combination fails. After the last phase, the orchestrator runs the plan’s global verification commands against the accumulated tree, then one reviewer pass over the whole change set — including every newly created file — against the goal, the success criteria, and the non-goals. Findings route back, affected checks re-run, and only then is it done.
Then one more thing, cheap because the material is already collected: if the
run recorded deviations or consciously declined findings, the orchestrator
closes with a one-line-each retro on the ones that carry durable repository
knowledge — a plan assumption the codebase contradicted, a missing test
wrapper, a convention no document states — and names where each belongs:
CLAUDE.md, a bin/ wrapper, the plan template. Suggestions only; the
orchestrator cannot create files, and the user decides what gets recorded. The
deviation log preserves what happened; the retro is for the subset that
should change what the next session reads before it starts.
Coverage lives in the reviewer, filtering lives in the orchestrator
The reviewer reports every plausible finding in its categories, tagged with a severity and a confidence, and filters nothing. The orchestrator decides what matters, routes those back, and records in the plan file anything it consciously declines.
This split exists because a severity filter inside the reviewer reads as an instruction to withhold, and current models follow it literally: they investigate just as hard, find the bug, and decline to mention it. Precision goes up, recall goes down, and real defects vanish silently. See Prompting.
The reviewer is also told not to write the fix — file, symbol, evidence, impact, smallest conceptual correction. Two reasons: an expensive model producing patch text is expensive output for work the cheap model is about to do anyway, and a pasted patch collapses the role split that makes the review independent.
What the orchestrator is actually prevented from doing
“The orchestrator writes no code” is two different guarantees, and only one of them is enforced.
Write is withheld through the skill’s disallowed-tools, so it cannot create a
file at all — no new component, no new test, no new migration appearing from the
orchestrator tier. That part is the harness’s job and it holds.
Edit it keeps, because the plan file needs editing for phase status and
deviations, and frontmatter cannot scope a tool to one path. So “the plan file
only” is a rule the orchestrator holds itself to. If you need that enforced
rather than instructed, a PreToolUse hook on Edit is the instrument — with
the caveat that hooks are dropped for plugin-packaged skills, so it protects only
the install.sh path.
An allowlist would be the stronger-looking option and is the wrong one here:
omitting ToolSearch makes the deferred SendMessage unloadable, which silently
breaks resuming the same implementer inside a phase.
Delegate, then commit to the delegation
An orchestrator that re-derives what a subagent already reported pays for the same work twice at the more expensive tier. The rule is: read the diff to verify the work, not to repeat it. Corrections inside a phase resume the same implementer (it keeps its context); the next phase gets a fresh one; reviewers are always fresh.
When not to use this
A fresh subagent starts cold: no conversation, no prior file reads, no warm
cache. So orchestrator → implementer → orchestrator reads the diff can pay
twice to read the same problem. For a three-file feature, one continuous Sonnet
session is often both cheaper and better oriented.
Isolation earns its cost when the implementation produces a lot of output, the run is long, or the phases are genuinely independent.
| Work | Workflow |
|---|---|
| Small, deterministic edit | Sonnet directly; batched Haiku for mechanical sets |
| Normal feature, clear multi-file bug | Sonnet directly, optionally one closing review |
| Long-running or context-heavy feature | /make-plan → /execute-plan |
| Migration, auth, security, public API, AI integration | Opus plan → /execute-plan → review gates |
| Genuinely open-ended architecture | Fable plan → same execution path |
| Plan turns out to be wrong | Stop, fix the plan with the big model, record the deviation, resume |
A typo, an obvious one-file bug, or a fully specified UI tweak does not need a plan file at all.
Stopping a stuck implementer
The tempting control is maxTurns, and it is the wrong one twice over. Even if
it worked it would be wrong at small values: a turn is every agentic step — read
the packet, grep, read the service, read the test, edit, run, read the failure,
fix — so a low cap truncates correct work mid-phase and leaves half-written code
behind. And it does not work; see below.
The failure actually worth stopping is edits that gain no information. So the implementer is told: if two materially different repair attempts produce the same failure and nothing was learned between them, stop and report the failing command, the hypotheses tested, what changed between attempts, the evidence, and the most likely next investigation. Three red tests in a row is often just a correct sequence — the reproduction test failing as designed, then an edge case surfacing, then a stale fixture.
Do not reach for maxTurns as the backstop either. Measured on a subagent that
declared maxTurns: 15: it ran 65 turns. The key is accepted and ignored, so
the instruction above is the whole mechanism — there is no hard limit behind it.
Why the read-only agents carry no turn limit
planning-researcher and plan-reviewer had maxTurns: 15 and it was removed.
The reasoning above applies to them with more force, not less.
The first reason is that it was not doing anything. A researcher delegated five
bugs spanning two repositories was reported back to the planner as
Done · 21 tool uses · 49.8k tokens · 1m 16s — and its subagent transcript
shows it was at turn 33 of an eventual 65 when that happened. A cap of 15
did not stop it at 15, or at all. Inert config is worse than no config: it reads
like a control, so you stop looking for the real one.
What the planner received as the agent’s “final answer” was its interleaved
narration — “Found AppSetting, EnforceAppSettings middleware… let’s dig
in” — because the run was cut mid-investigation, before the turn that writes
the report. The planner said the answer kept truncating and went back to
grepping the codebase itself at planner rates. Its follow-up then resumed the
agent in the background, where it ran another 32 turns that nothing was waiting
for.
So the shape of the failure is real even though the cap was not its cause: an
agent that searches first and reports last loses everything if it is stopped
early, and the visible symptom is not “stopped early” but “the agent lost its
answer”. The defence is to size the delegation so it finishes — see the research
section of make-plan — not to add a second stopping condition.
A cap is also the wrong place to economise. Research is already the cheap tier —
planning-researcher runs Sonnet at medium precisely so breadth-first
grepping does not run at planner rates. That is where the saving comes from. A
turn cap on top of it buys nothing and buys it at the price of the report.
The risk a cap guards against is a runaway loop, and what makes a loop
expensive is repeated writes. Neither agent has Edit or Write. Both keep
Bash, because breadth-first grep and git log are the job — a measured
research run made thirty-odd Bash calls and a handful of Reads. Bash is not a
sandbox, and permissionMode: plan gates it only in a local install; a plugin
install drops that key (see gotchas.md).
So the guarantee here is narrow, and worth stating exactly rather than rounding up to “read-only”: these agents will not sit in a write-test-write loop, so a search that runs ten turns too long costs ten cheap turns and nothing else. A truncated report costs the whole delegation.
Cap agents that edit. Let agents that only look, finish. And before trusting any
cap, check a subagent transcript under
~/.claude/projects/<project>/<session>/subagents/ and count the turns it
actually took.
Effort
effort is available on both skills and agents: low, medium, high,
xhigh, max, or an integer. The top levels are model-gated.
The split that matters is orchestrator versus implementer. Finding the next
phase, assembling a packet, and checking a box do not need the same reasoning
budget as writing the code. So the orchestrator skill runs medium and
implementer overrides back up to high. Without the explicit effort: high on
the implementer it would inherit the orchestrator’s medium.
The orchestrator’s medium is conditional, and the condition is invisible while
it holds. Skill frontmatter applies for the current turn only, so medium
survives a run only while every delegation stays in the foreground: a
backgrounded worker’s completion notification starts a new turn at the
session’s effort. Measured medium→xhigh at the first notification in three
runs (2.1.239, 2.1.251, 2.1.259). Two defences, both in the settings snippet —
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 in the project env keeps delegations
in the foreground, and the session should be set to Sonnet at medium before
/execute-plan is invoked (modelSettings.claude-sonnet-5.effortLevel makes
that standing). Both are overridden by CLAUDE_CODE_EFFORT_LEVEL if it is set.
See Gotchas.
Agent effort is fixed per definition — there is no per-invocation knob. So
“bump the final review to xhigh” is not something the orchestrator can do; it
would need a second agent definition. The reviewer stays at high, and the final
gate is deepened by telling it in the delegation that this is the whole
accumulated change set against every success criterion.