plan → execute

Gotchas

Things that break this setup silently — where it keeps running and quietly stops doing what you think it does.

The checkable ones are mechanized: scripts/doctor.sh <project> asserts them against a live install — the env override, the dual distribution paths, a dropped permissionMode, inert maxTurns keys, unadapted worker agents, wrapper refusal (--probe), a stale plugin snapshot. Prose is for understanding a failure; run the doctor to find one.

Do not run /make-plan from plan mode

Plan mode injects its own workflow into the session, and that injection outranks a skill body. Two things follow, both observed in a real run:

The symptom is a plan in the wrong directory and no planning-researcher anywhere in the agent list. Hardening the skill’s wording does not fix it; the skill never had the last word. /make-plan therefore checks for plan mode and refuses, and plansDirectory is worth setting so that even the plan-mode path stays inside the repository.

message.model is not the pin record, and what it records changed

The authoritative record of a skill’s model pin is the command_permissions attachment emitted when the skill is invoked:

{"type": "command_permissions", "allowedTools": ["Read","Glob","Grep","Bash","Write","Agent"], "model": "claude-opus-5"}
{"type": "command_permissions", "allowedTools": [], "model": "claude-sonnet-5"}

A skill with no model: pin logs no model key at all — so the key’s presence is itself the signal that frontmatter was applied. scripts/verify-models.py pulls this out of a transcript for you.

The model field on assistant records answers a different question, and the answer changed between versions. On 2.1.220–2.1.239 it recorded the session’s configured model for the whole run: one /execute-plan session logged claude-opus-5 on all 521 main-thread turns while the skill was executing on Sonnet. On 2.1.259 it recorded the effective model: an Opus session logged claude-sonnet-5 for the /execute-plan turn, then claude-opus-5 from the first subagent notification — which is where the pin was lost, not where the tiering was designed to hand over. Either way, read as “which model ran this skill” it gives a wrong answer, for two different reasons.

The field that held up across every version measured here is the top-level effort on assistant records: it is present on every one of them, and it moved at the pin drop in all three runs below.

For subagent transcripts, under <session>/subagents/agent-<id>.jsonl, message.model is worth reading: a worker’s “session” is its own run, so the model recorded there is what the worker actually ran — the frontmatter pin, unless the delegation passed an explicit model override. Measured on 2.1.238, one Sonnet-configured /execute-plan session: the implementer’s transcript records claude-sonnet-5, the quick-implementer’s records Haiku, the reviewer’s records claude-opus-5 — three tiers, all diverging from nothing, each matching its pin. The adjacent agent-<id>.meta.json names the agentType, which is what lets verify-models.py attribute token usage per tier.

A skill’s model pin dies at the first subagent notification

A skill’s model: and effort: frontmatter “applies for the rest of the current turn”, and “the session model resumes when you send your next prompt” (skills docs). The second half is the trap: a subagent completion is a new turn, and nobody sent a prompt.

Every delegation is backgrounded by default now, and its completion arrives as a user record with origin.kind: task-notification and promptSource: system. That record starts a new turn, and the new turn runs on the session model at the session effort. Nothing in the transcript announces it.

Measured in three real /execute-plan runs in one project, with no user input between the skill turn and the flip:

Date Version Session Skill turn From the first notification
2026-08-22 2.1.239 Fable, xhigh fable, medium fable, xhigh (12:09)
2026-08-29 2.1.251 Sonnet, xhigh sonnet, medium sonnet, xhigh
2026-09-06 2.1.259 Opus, xhigh claude-sonnet-5, medium claude-opus-5, xhigh (17:28:44)

The 2.1.259 run then cascaded into the worker tier: the now-Opus orchestrator passed model: opus explicitly on all seven implementer delegations — the Agent tool’s per-invocation model parameter overrides the agent file’s model: — and those subagent transcripts record claude-opus-5 against a model: sonnet agent definition. One dropped pin can take the rest of the tiering with it.

/make-plan has the same shape from the other side: a 2026-08-18 run in a Fable session, skill pinned opus, ran everything after the first planning-researcher notification on Fable at high — including writing the plan.

Three defences, and use all of them:

The one-line check: the delegation’s tool result is the worker’s report rather than “Async agent launched successfully”, verify-models.py counts zero task-notifications, and the effort field on assistant records stays at the skill’s value for the whole run. verify-models.py prints a WARNING naming the exact record where it stopped.

CLAUDE_CODE_SUBAGENT_MODEL overrides every model: pin

It takes first precedence over agent frontmatter. Set it to Sonnet “for consistency” and you have silently overridden the Haiku quick-implementer and the Opus plan-reviewer — the tiering is gone and nothing tells you. Do not set it. Frontmatter pins are targeted and sufficient.

Managed settings can do the same thing through availableModels and modelOverrides, which remap aliases at the policy level.

Plugin agents drop permissionMode, hooks, and mcpServers

Claude Code says so directly: “sets <field>, which is ignored for plugin agents. Use .claude/agents/ for this level of control.”

planning-researcher and plan-reviewer use permissionMode: plan to be unable to write. Installed via the plugin they are only instructed not to. If that distinction matters, use install.sh.

Do not run both distribution paths at once

Plugin agents are namespaced — plan-and-execute:planning-researcher — so they do not shadow a project’s .claude/agents/planning-researcher.md. Both load, under different names, and the orchestrator picks one. Measured here: it picked the plugin copy, the one that had lost permissionMode: plan, while the enforced project copy sat unused. The read-only guarantee degraded to prompt-only and nothing said so.

Pick one path per project. If projects are installed with install.sh, leave the plugin disabled in ~/.claude/settings.json.

/plugin marketplace add /path/to/repo copies the tree into ~/.claude/plugins/cache/ at install time. Editing the repo afterwards changes nothing that Claude Code loads, and the cache stays pinned to the installed version string — so bumping only file contents, without bumping version in plugin.json, gives you no way to tell a stale install from a fresh one. Develop against install.sh; use the plugin to test the packaged path.

Built-in Plan and Explore skip CLAUDE.md

And built-in Plan inherits the main-session model — so in a planning session it runs your expensive planner to do grep work. That is why planning-researcher exists: it is pinned to Sonnet regardless of the session model, and, being a custom agent, it does inherit CLAUDE.md.

Auto-delegation not happening? It is the description

Almost always. Start it with “Use this agent to…” and include “use proactively”. The description is what the orchestrator matches against; the body is only read after the agent is chosen.

Model aliases float

model: opus became a different model on a release, with the same config and no warning. Capability went up; behaviour moved too, and a prompt tuned to the previous model can get quietly worse — see Prompting for the severity-filter case, which is exactly this.

Pin a full model ID only if you actually want frozen behaviour. You then own the upgrade manually, and you lose portability: aliases resolve differently across providers, and on some setups sonnet still points at an older generation. You cannot have both provider portability and a guaranteed model version — pick one deliberately.

plansDirectory must be inside the project root

“Custom directory for plan files, relative to project root. If not set, defaults to ~/.claude/plans/.” An absolute path outside the repository is rejected.

Subagent resume is real, and reviewers must not use it

A parent session can continue a completed subagent by messaging its agent id; the subagent keeps its full prior context. Use it for corrections inside the current phase — the implementer already knows what it did. Never for a reviewer: the value of a review is the fresh context, and a resumed reviewer is reviewing its own framing.

Subagents run in the background since 2.1.232, and run_in_background is gone

Both skills here are built on the opposite assumption: the planner needs the research before it can write the plan, and the orchestrator needs the code on disk before step 4 can diff it. Backgrounded, the orchestrator reviews an empty diff, the planner starts re-deriving evidence that is still in flight, and both lose their model pin at the notification (A skill’s model pin dies at the first subagent notification, above).

run_in_background: false used to be the instruction for this, and it is dead text now. Across 355 Agent calls in local transcripts (2.1.233–2.1.272) it was passed 5 times, and all 5 tool results came back “Async agent launched successfully.” — the harness ignored it. The current tool schema, on 2.1.272, has no such parameter at all.

The sub-agents docs say why. Fork mode is on by default in interactive sessions from 2.1.232, and with it on, Claude Code “runs the subagent in the background, forks and non-fork subagents alike, and Claude can’t ask for the foreground” — the run_in_background parameter is removed from the tool. Every measured pin drop above is on a version past that line. The agent frontmatter field background only has a documented true meaning (keep this agent in the background even when Claude asks for the foreground); background: false is not a foreground request, and an earlier draft of this fix that relied on it was wrong.

Two documented ways out, and only the first is deterministic:

Not measured here. The switch is documented; an interactive run with it set had not been recorded when this was written. A headless claude -p run is not evidence either way: fork mode is off there by default, so delegations return in the foreground with or without any setting — which is exactly how the background: false draft looked verified when it was not.

The symptoms of losing this do not look like a scheduling bug. They look like the agent losing its answer, and like a run that silently changed model.

maxTurns on an agent is not enforced — verify before trusting it

planning-researcher carried maxTurns: 15. A run of it was measured at 65 assistant turns. The key is accepted in frontmatter and does nothing, which is worse than omitting it: it reads like a control, so you stop looking for the one that actually stopped your agent.

Count for yourself rather than believing the frontmatter — subagent transcripts are written to ~/.claude/projects/<project>/<session>/subagents/agent-<id>.jsonl, one JSON record per line, and type == "assistant" is a turn.

It is gone from both read-only agents. See “Why the read-only agents carry no turn limit” in design.md.

A cut-off agent returns its narration as its answer

When a research or review agent is stopped mid-run — for whatever reason — what comes back is not an error. It is whatever text the agent had emitted so far, which for these agents is the running commentary between tool calls: “Found the middleware, let’s dig in.” It reads like a report that trails off.

Measured: a delegation reported back as Done · 21 tool uses · 49.8k tokens · 1m 16s had, at that moment, written no report at all — and the parent’s follow-up resumed it in the background for another 32 turns whose output nothing was waiting for.

So “the agent lost its answer” is usually “the agent never got to write one”. Size delegations so they finish, and treat a report that ends on a “now let’s check…” as a truncation, not a finding.

effort is fixed per agent definition

There is no per-invocation effort knob. “Bump the final review to xhigh” is not something the orchestrator can do — it would need a second agent definition. If you want a deeper final gate, either add one, or deepen it in the delegation prompt by naming the scope: the whole accumulated change set, every success criterion.

opusplan is the built-in near-equivalent

The built-in opusplan mode has the big model plan and a cheaper one execute, which approximates sessions 1 and 2 for medium-complexity work in one command. The manual split still wins when you want the plan on disk: opusplan keeps it in context, so it dies with the session.

Fast mode buys latency, not intelligence

/fast on Opus is real Opus with faster output, not a downgrade to a smaller model — but it is priced as a premium. Worth it for an orchestrator you are watching; wasteful for a background review gate.

Duplication budget

CLAUDE.md holds universal invariants. Agent files hold the role contract plus at most the one or two rules that bite that role hardest. The skill holds the orchestration procedure.

Repeating a long pinned test command in both worker agents is a symptom, not a solution — put it behind a bin/test-<lang> wrapper and the duplication shrinks to one line everywhere. See templates/bin/test-example.sh.