Gotchas
Things that break this setup silently — where it keeps running and quietly stops doing what you think it does.
The checkable ones are mechanized:
scripts/doctor.sh <project>
asserts them against a live install — the env override, the dual distribution
paths, a dropped permissionMode, inert maxTurns keys, unadapted worker
agents, wrapper refusal (--probe), a stale plugin snapshot. Prose is for
understanding a failure; run the doctor to find one.
Do not run /make-plan from plan mode
Plan mode injects its own workflow into the session, and that injection outranks a skill body. Two things follow, both observed in a real run:
- “In this phase you should only use the Explore subagent type” — so every
research delegation goes to
Explore, which inherits the session model. The cheap researcher tier is bypassed and breadth-first grepping runs at planner rates. - “this is the only file you are allowed to edit”, pointing at the harness’s
own plan path. The plan lands in
~/.claude/plans/<auto-slug>.mdinstead ofplans/<PlanName>.md, outside version control, usually followed by a copy — so now there are two.
The symptom is a plan in the wrong directory and no planning-researcher
anywhere in the agent list. Hardening the skill’s wording does not fix it; the
skill never had the last word. /make-plan therefore checks for plan mode and
refuses, and plansDirectory is worth setting so that even the plan-mode path
stays inside the repository.
message.model is not the pin record, and what it records changed
The authoritative record of a skill’s model pin is the command_permissions
attachment emitted when the skill is invoked:
{"type": "command_permissions", "allowedTools": ["Read","Glob","Grep","Bash","Write","Agent"], "model": "claude-opus-5"}
{"type": "command_permissions", "allowedTools": [], "model": "claude-sonnet-5"}
A skill with no model: pin logs no model key at all — so the key’s
presence is itself the signal that frontmatter was applied.
scripts/verify-models.py
pulls this out of a transcript for you.
The model field on assistant records answers a different question, and the
answer changed between versions. On 2.1.220–2.1.239 it recorded the
session’s configured model for the whole run: one /execute-plan session
logged claude-opus-5 on all 521 main-thread turns while the skill was
executing on Sonnet. On 2.1.259 it recorded the effective model: an Opus
session logged claude-sonnet-5 for the /execute-plan turn, then
claude-opus-5 from the first subagent notification — which is where the pin
was lost, not where the tiering was designed to hand over. Either way, read as
“which model ran this skill” it gives a wrong answer, for two different reasons.
The field that held up across every version measured here is the top-level
effort on assistant records: it is present on every one of them, and it
moved at the pin drop in all three runs below.
For subagent transcripts, under <session>/subagents/agent-<id>.jsonl,
message.model is worth reading: a worker’s “session” is its own run, so the
model recorded there is what the worker actually ran — the frontmatter pin,
unless the delegation passed an explicit model override. Measured on 2.1.238, one
Sonnet-configured /execute-plan session: the implementer’s transcript records
claude-sonnet-5, the quick-implementer’s records Haiku, the reviewer’s records
claude-opus-5 — three tiers, all diverging from nothing, each matching its
pin. The adjacent agent-<id>.meta.json names the agentType, which is what
lets verify-models.py attribute token usage per tier.
A skill’s model pin dies at the first subagent notification
A skill’s model: and effort: frontmatter “applies for the rest of the
current turn”, and “the session model resumes when you send your next
prompt” (skills docs). The second
half is the trap: a subagent completion is a new turn, and nobody sent a prompt.
Every delegation is backgrounded by default now, and its completion arrives as a
user record with origin.kind: task-notification and promptSource: system.
That record starts a new turn, and the new turn runs on the session model at
the session effort. Nothing in the transcript announces it.
Measured in three real /execute-plan runs in one project, with no user input
between the skill turn and the flip:
| Date | Version | Session | Skill turn | From the first notification |
|---|---|---|---|---|
| 2026-08-22 | 2.1.239 | Fable, xhigh |
fable, medium |
fable, xhigh (12:09) |
| 2026-08-29 | 2.1.251 | Sonnet, xhigh |
sonnet, medium |
sonnet, xhigh |
| 2026-09-06 | 2.1.259 | Opus, xhigh |
claude-sonnet-5, medium |
claude-opus-5, xhigh (17:28:44) |
The 2.1.259 run then cascaded into the worker tier: the now-Opus orchestrator
passed model: opus explicitly on all seven implementer delegations — the
Agent tool’s per-invocation model parameter overrides the agent file’s
model: — and those subagent transcripts record claude-opus-5 against a
model: sonnet agent definition. One dropped pin can take the rest of the
tiering with it.
/make-plan has the same shape from the other side: a 2026-08-18 run in a Fable
session, skill pinned opus, ran everything after the first
planning-researcher notification on Fable at high — including writing the
plan.
Three defences, and use all of them:
- Keep the delegations in the foreground. Set
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1in the project’s.claude/settings.jsonenv, so a delegation’s report comes back as its tool result and the turn never ends. See Subagents run in the background since 2.1.232 below for why nothing in an agent file can do this. - Leave
CLAUDE_CODE_EFFORT_LEVELunset. It takes precedence over/effort,--effort,modelSettings,effortLeveland skill or subagent frontmatter (env vars, model config). Set toxhigh, it silently overrides the orchestrator’smediumand the researcher’smediumwhether or not the pin holds. - Set the session, not just the skill.
/model sonnetand/effort mediumbefore/execute-plan, orclaude --model sonnet --effort medium./clearkeeps the session’s current model and effort, so clearing is not setting them.modelSettings.<model>.effortLevelin settings is the standing version of the same safety net — seetemplates/settings.snippet.json.
The one-line check: the delegation’s tool result is the worker’s report rather
than “Async agent launched successfully”, verify-models.py counts zero
task-notifications, and the effort field on assistant records stays at the
skill’s value for the whole run. verify-models.py prints
a WARNING naming the exact record where it stopped.
CLAUDE_CODE_SUBAGENT_MODEL overrides every model: pin
It takes first precedence over agent frontmatter. Set it to Sonnet “for
consistency” and you have silently overridden the Haiku quick-implementer and
the Opus plan-reviewer — the tiering is gone and nothing tells you. Do not set
it. Frontmatter pins are targeted and sufficient.
Managed settings can do the same thing through availableModels and
modelOverrides, which remap aliases at the policy level.
Plugin agents drop permissionMode, hooks, and mcpServers
Claude Code says so directly: “sets <field>, which is ignored for plugin
agents. Use .claude/agents/ for this level of control.”
planning-researcher and plan-reviewer use permissionMode: plan to be
unable to write. Installed via the plugin they are only instructed not to.
If that distinction matters, use install.sh.
Do not run both distribution paths at once
Plugin agents are namespaced — plan-and-execute:planning-researcher — so they
do not shadow a project’s .claude/agents/planning-researcher.md. Both load,
under different names, and the orchestrator picks one. Measured here: it picked
the plugin copy, the one that had lost permissionMode: plan, while the enforced
project copy sat unused. The read-only guarantee degraded to prompt-only and
nothing said so.
Pick one path per project. If projects are installed with install.sh, leave the
plugin disabled in ~/.claude/settings.json.
A directory-source plugin is a snapshot, not a live link
/plugin marketplace add /path/to/repo copies the tree into
~/.claude/plugins/cache/ at install time. Editing the repo afterwards changes
nothing that Claude Code loads, and the cache stays pinned to the installed
version string — so bumping only file contents, without bumping version in
plugin.json, gives you no way to tell a stale install from a fresh one. Develop
against install.sh; use the plugin to test the packaged path.
Built-in Plan and Explore skip CLAUDE.md
And built-in Plan inherits the main-session model — so in a planning session it
runs your expensive planner to do grep work. That is why
planning-researcher exists: it is pinned to Sonnet regardless of the session
model, and, being a custom agent, it does inherit CLAUDE.md.
Auto-delegation not happening? It is the description
Almost always. Start it with “Use this agent to…” and include “use proactively”.
The description is what the orchestrator matches against; the body is only read
after the agent is chosen.
Model aliases float
model: opus became a different model on a release, with the same config and no
warning. Capability went up; behaviour moved too, and a prompt tuned to the
previous model can get quietly worse — see Prompting for the
severity-filter case, which is exactly this.
Pin a full model ID only if you actually want frozen behaviour. You then own the
upgrade manually, and you lose portability: aliases resolve differently across
providers, and on some setups sonnet still points at an older generation. You
cannot have both provider portability and a guaranteed model version — pick one
deliberately.
plansDirectory must be inside the project root
“Custom directory for plan files, relative to project root. If not set, defaults
to ~/.claude/plans/.” An absolute path outside the repository is rejected.
Subagent resume is real, and reviewers must not use it
A parent session can continue a completed subagent by messaging its agent id; the subagent keeps its full prior context. Use it for corrections inside the current phase — the implementer already knows what it did. Never for a reviewer: the value of a review is the fresh context, and a resumed reviewer is reviewing its own framing.
Subagents run in the background since 2.1.232, and run_in_background is gone
Both skills here are built on the opposite assumption: the planner needs the research before it can write the plan, and the orchestrator needs the code on disk before step 4 can diff it. Backgrounded, the orchestrator reviews an empty diff, the planner starts re-deriving evidence that is still in flight, and both lose their model pin at the notification (A skill’s model pin dies at the first subagent notification, above).
run_in_background: false used to be the instruction for this, and it is dead
text now. Across 355 Agent calls in local transcripts (2.1.233–2.1.272) it was
passed 5 times, and all 5 tool results came back “Async agent launched
successfully.” — the harness ignored it. The current tool schema, on 2.1.272,
has no such parameter at all.
The sub-agents docs say why.
Fork mode is on by default in interactive sessions from 2.1.232, and with it
on, Claude Code “runs the subagent in the background, forks and non-fork
subagents alike, and Claude can’t ask for the foreground” — the
run_in_background parameter is removed from the tool. Every measured pin drop
above is on a version past that line. The agent frontmatter field background
only has a documented true meaning (keep this agent in the background even
when Claude asks for the foreground); background: false is not a foreground
request, and an earlier draft of this fix that relied on it was wrong.
Two documented ways out, and only the first is deterministic:
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1— “runs the subagent in the foreground, in every kind of session and whether or not fork mode is on”. Set it in the project’s.claude/settings.jsonenv(templates/settings.snippet.json);scripts/doctor.shasserts it. The cost is that it also disablesrun_in_backgroundon Bash and the Ctrl+B shortcut in that project.CLAUDE_CODE_FORK_SUBAGENT=0turns fork mode off, after which Claude runs a subagent “in the foreground when it needs the result before continuing” — Claude’s judgment, not a guarantee.
Not measured here. The switch is documented; an interactive run with it set
had not been recorded when this was written. A headless claude -p run is not
evidence either way: fork mode is off there by default, so delegations return
in the foreground with or without any setting — which is exactly how the
background: false draft looked verified when it was not.
The symptoms of losing this do not look like a scheduling bug. They look like the agent losing its answer, and like a run that silently changed model.
maxTurns on an agent is not enforced — verify before trusting it
planning-researcher carried maxTurns: 15. A run of it was measured at 65
assistant turns. The key is accepted in frontmatter and does nothing, which is
worse than omitting it: it reads like a control, so you stop looking for the one
that actually stopped your agent.
Count for yourself rather than believing the frontmatter — subagent transcripts
are written to
~/.claude/projects/<project>/<session>/subagents/agent-<id>.jsonl, one JSON
record per line, and type == "assistant" is a turn.
It is gone from both read-only agents. See “Why the read-only agents carry no
turn limit” in design.md.
A cut-off agent returns its narration as its answer
When a research or review agent is stopped mid-run — for whatever reason — what comes back is not an error. It is whatever text the agent had emitted so far, which for these agents is the running commentary between tool calls: “Found the middleware, let’s dig in.” It reads like a report that trails off.
Measured: a delegation reported back as Done · 21 tool uses · 49.8k tokens ·
1m 16s had, at that moment, written no report at all — and the parent’s
follow-up resumed it in the background for another 32 turns whose output nothing
was waiting for.
So “the agent lost its answer” is usually “the agent never got to write one”. Size delegations so they finish, and treat a report that ends on a “now let’s check…” as a truncation, not a finding.
effort is fixed per agent definition
There is no per-invocation effort knob. “Bump the final review to xhigh” is not
something the orchestrator can do — it would need a second agent definition. If
you want a deeper final gate, either add one, or deepen it in the delegation
prompt by naming the scope: the whole accumulated change set, every success
criterion.
opusplan is the built-in near-equivalent
The built-in opusplan mode has the big model plan and a cheaper one execute,
which approximates sessions 1 and 2 for medium-complexity work in one command.
The manual split still wins when you want the plan on disk: opusplan keeps
it in context, so it dies with the session.
Fast mode buys latency, not intelligence
/fast on Opus is real Opus with faster output, not a downgrade to a smaller
model — but it is priced as a premium. Worth it for an orchestrator you are
watching; wasteful for a background review gate.
Duplication budget
CLAUDE.md holds universal invariants. Agent files hold the role contract plus
at most the one or two rules that bite that role hardest. The skill holds the
orchestration procedure.
Repeating a long pinned test command in both worker agents is a symptom, not a
solution — put it behind a bin/test-<lang> wrapper and the duplication shrinks
to one line everywhere. See
templates/bin/test-example.sh.