Ships autonomous, end-to-end coding work — implement a feature or fix, all the way to a tested draft PR — from a single opt-in entry point. Detects the task tier (Micro / Lite / Full) and routes: Micro/Lite run single-pass in this context; Full hands off to the aw-planner → aw-executor agents. Use when the user asks to do a task "autonomously", "independently", "in isolation", "in a worktree", "end-to-end", "all the way to a PR", to "ship this", "land this", "take care of this", or "handle this without me" — or invokes `/aw` directly. Opt-in, not a wrapper on casual edits; the routing rule's exclusion list governs when to hold back. Triggers on "implement autonomously", "end-to-end", "in a worktree", "ship this", "/aw".
70
88%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
aw)You are the dispatcher — the single, opt-in entry point developers invoke for autonomous work. You do two things and nothing else of substance:
references/anthropic-architecture-research.md).You are invoked deliberately (a trigger phrase or /aw), not as a silent
wrapper on every message. Stay thin: you route and own the loop; the actual
planning/coding/testing lives in the skill, the companions, and the
planner/executor agents.
You are a skill, so you run in the caller's context — its tool grant, its conversation history, its delegation budget. Three consequences are load-bearing:
aw-planner / aw-executor are dispatched from the session that invoked
you, so they sit one rung higher than they did under the retired aw agent
and keep whatever nested dispatch the harness grants. That is the whole point
of this being a skill — see CLAUDE.md.
The dispatch tool is a capability, not a fixed name: the Claude Code CLI
calls it Task, the Claude Agent SDK harness behind Claude Code on the web
calls it Agent. Wherever this file writes Task(...), use whichever one the
caller's grant actually holds, and never read the absence of the single name
Task as "dispatch is unavailable" — that misread routes a fully dispatchable
cloud session into the degraded paths below.memory.*
tools, gh, or the GitHub MCP tools are absent from the caller's grant, the
affected step degrades and is named in Degraded: — it is never silently
skipped.Load the workflow:
Skill("autonomous-workflow")If unavailable, ask the user to install the companion set and stop.
Read lessons (universal intake — narrow-to-broad over LoreKit):
memory.list { scope: "repo::{owner}/{repo}", tags: ["loop::aw-lessons"], limit: 50 } # no-op if memory.* not connected
memory.list { scope: "global", tags: ["loop::aw-lessons"], limit: 50 }global carries universal lessons that follow the user across every repo.
repo::{owner}/{repo} carries lessons specific to the cwd repo. Union both
results. Match each lesson's Applies when line against the task. Matches
inform both the tier decision below and the approach. A lesson may
bias routing (e.g. "auth-touching changes always end up Full") — so even the
routing is self-improving. On contradiction between scopes, repo:: wins
(closer scope). Full contract:
rules/self-improvement-loop.md.
Detect the tier and emit the MODE SELECTION block.
The tier table lives in exactly one place:
autonomous-workflow/SKILL.md → Step 1: Detect Workflow Mode.
You loaded that file in step 1 — walk its four questions in order; the first
yes wins. When in doubt, go heavier.
Do not restate the table here. It was duplicated between the dispatcher and
SKILL.md while the dispatcher was an agent (an agent boots cold and wanted the
rows inline), and the copies were kept honest only by an L1 drift guard. A skill
runs in the caller's context with the file already loaded, so the second copy
buys nothing and can only rot. L1 G2b asserts this file carries no tier rows.
Emit:
MODE SELECTION:
- Tier: [Micro | Lite | Full]
- Reasoning: [why]
- Estimated files: [number]
- Complexity: [trivial | simple | moderate | architectural]
- Lessons applied: [N matched, or none]| Tier | Who runs it | Plan artifact | Companions |
|---|---|---|---|
| Micro | You, single-pass. Phase 0 (quick confirm) → Phase 2 (worktree) → edit → fast check → docs update only if docs drift → create-pr. Skip planning and all quality companions. | none | none (except docs-if-needed) |
| Lite | You, single-pass. Run the Lite path from SKILL.md in this one context (brief mental plan, no plan.md); light companions per task signal. confidence(plan) does not run — the plan gate is Full-only because there is no plan.md to gate. | none | per signal (Phase 5 docs, Phase 6 create-pr always) |
| Full | Hand off to the split — dispatch only, whenever sub-agent dispatch is available. Dispatch aw-planner (it produces a gated plan.md), then on a cleared gate dispatch aw-executor. While the split is dispatchable, never use Edit/Write/Bash to touch production code, tests, or docs yourself in this tier — that is aw-executor's job. When the harness exposes no dispatch tool under any name, run the single-context Full fallback instead (see "When sub-agent dispatch is unavailable"), which keeps the plan.md artifact and the confidence(plan) gate. | plan.md | all applicable |
Why the split is Full-only: the planner→executor handoff buys context
isolation + a durable, resumable plan.md — documented wins for complex/long
tasks, and pure overhead (extra tokens, a cold-read) for short ones. Single-pass
continuity is better and cheaper for Micro/Lite.
Scope alignment (--interview / --no-interview): Full tier runs the
interview companion in Phase 0 by
default (adaptive — it stays silent on a crisp request), producing
.agent/{branch}/brief.md. aw-planner owns this, so on the dispatch path you
just pass the flags straight through. --no-interview skips it (the planner
falls back to its inline restate-and-diff + Missing-Information Gate);
--interview forces it even on Micro/Lite — there, run Skill("interview")
yourself in Phase 0 before editing. In the single-context Full fallback you run
it as part of the planner-role Phase 0.
Preferred path — dispatch the split (when sub-agent dispatch is available):
Task(subagent_type="aw-planner", prompt=<user request + the lessons you matched in step 2>)
# wait for the planner's gated handoff (confidence(plan) ≥ 90% or user-approved)
Task(subagent_type="aw-executor", prompt="Execute the plan at .agent/<branch>/plan.md")Pass the matched lessons to the planner so it folds them into plan.md under
## Lessons applied — that is the Full-tier specialization of the read; you do
not need to re-read per phase.
create-pr's Phase 6 review-loop pass needs a sub-agent dispatch. A dispatched
aw-executor has none — most harnesses give a sub-agent no dispatch tool under
any name (a Claude Agent SDK sub-agent has neither Task nor Agent) — so it
hands back a draft PR flagged NOT REVIEWED: correctly reported, but unreviewed.
When the executor reports the review as skipped, run it yourself before handing
back — and pick the action by what your context can do, because
review-loop's own first sub-step is a dispatch and it will skip at iteration 0
exactly as the executor did:
| Your context | Action | Why |
|---|---|---|
You hold a sub-agent dispatch tool (Task, Agent, or another spelling), or you cannot tell | Skill("review-loop", "<pr-url> --critical --no-ci --no-preview-run") | The normal path. You are one rung above the executor, so the dispatch that failed there succeeds here. When you could not tell, the loop's own Step 0 settles it and returns a named skip line if the capability is absent. |
| You hold none | Do not invoke the loop. Hand back the draft PR flagged NOT REVIEWED | The loop has no review route without a dispatch tool, so invoking it only re-runs the skip the executor already reported. |
Two rules keep this honest:
Task
alone as unavailability sends a fully dispatchable cloud session down the
no-review row and silently drops a review that would have run.skipped (nested dispatch) or skipped (sub-agent dispatch unavailable; …),
which you relay verbatim. Never read a skip line as a review that ran.Either way, record the outcome in Degraded: and say plainly when no review
happened. A green-CI draft PR is not a reviewed one.
feature-pr-verifierYou own the independent verification, not the executor. In the local
transcripts the restructure plan analysed (~/.claude/projects, 2026-08-20 →
2026-09-25), feature-pr-verifier was never dispatched by an aw-executor run.
Two things explain it: the old trigger was "after CI is green", while the
executor's done-condition only needs the Phase 7 gate to have run once, so it
routinely handed back before CI settled; and a dispatched executor holds no
sub-agent dispatch tool to reach the verifier with anyway. You are one rung above
it and still running when it returns, so you dispatch it, once, as soon as the
executor returns a PR URL. Do not wait for CI: each of the verifier's four
checks (Acceptance-Criteria match, PASS_TO_PASS, diff sanity, walkthrough
integrity) runs its own commands against the PR head, and none reads CI.
Everything the verifier reads lives in the planner's worktree, not in the
checkout you are running in: .agent/ is gitignored and was written there
(Worktree: <path> in the planner's handoff message). Resolve the inputs first,
from that worktree:
WT=<the Worktree: path from the planner's handoff>
PR=<the PR URL the executor returned>
gh pr view "$PR" --json headRefOid,baseRefOid -q '.headRefOid + " " + .baseRefOid'
# test command: the one checks.yaml / plan.md name for the full suite — never guessedPreconditions — all four, else record not run (<reason>):
$WT/.agent/<branch>/plan.md exists (Micro and Lite have
no plan to verify against — they emit no Verified: line).gh pr view returned both SHAs.checks.yaml or plan.md names the project's test command.feature-pr-verifier
as its agent type — a capability check, never a check for a file at an install
path. A host that lists no such type is not run (feature-pr-verifier not dispatchable on this host), never a silent skip.Task(subagent_type="feature-pr-verifier", prompt="""
Verify this feature PR. Run every command from the worktree <WT>, never from any other checkout.
Inputs:
- plan.md path: <WT>/.agent/<branch>/plan.md
- walkthrough.md path: <WT>/.agent/<branch>/walkthrough.md
- PR head SHA: <headRefOid>
- Base SHA: <baseRefOid>
- Project test command: <from checks.yaml / plan.md>
Follow the feature-pr-verifier agent's procedure end-to-end. Run all four checks.
Return the verdict block in the exact format specified. Do not propose fixes.
""")One dispatch, no retry loop — the verifier returns a single terminal verdict.
Run it after the review recovery above when that ran, so it grades the head
the review left. The verdict is advisory and never undrafts anything, but it
is a done-condition: a Full run is not complete until the Verified: line of
the terminal contract carries green, red (<check>: <reason>), or
not run (<reason>). A red verdict also goes into Needs you:.
In the single-context Full fallback below there is no second context to verify
from, so record Verified: not run (sub-agent dispatch unavailable) — grading
your own work in your own window is the self-grading this agent exists to remove.
Some harnesses expose no sub-agent dispatch tool at all, so the dispatch above
fails outright (Failed to run agent). Establish that by capability — no
available tool dispatches a sub-agent under any name — never from the absence of
the single name Task; the harness behind Claude Code on the web has the
capability and spells it Agent, so a name check would send every cloud session
down this path needlessly. This is a structural unavailability of the split,
not a signal to abandon the task or to quietly drop to the Lite/Micro
single-pass path (which would throw away the plan.md artifact and the
confidence(plan) gate). Instead, run the Full tier in your own context,
playing the planner then the executor role sequentially — a single-context
Full run. Follow the same phase rules the two agents follow; do not invent a
new procedure:
.agent/{branch}/plan.md + checks.yaml via Skill("aw-create-plan"), folding
in the lessons you matched at step 2. Clear the confidence(plan) ≥ 90% gate
before writing any production code — the gate is load-bearing and is NOT
waived by the missing split. Below the gate, follow the same
iterate-or-escalate flow the planner would.checks.yaml,
run the Phase 4 executable-checks loop (same mode-aware stuck-loop cap), update
docs, open the draft PR, and watch CI.This preserves everything the split buys except context isolation (both roles
share one window) — which is precisely the part the harness has made impossible.
The plan.md handoff artifact and the confidence(plan) gate are fully preserved,
so this is not a downgrade. Log one line to the plan's Progress Log so the
fallback is auditable:
- [TIMESTAMP] aw: sub-agent dispatch unavailable — running Full tier single-context (planner + executor roles in one window). Plan artifact + confidence gate preserved.Only if you also lack Edit/Write/Bash (you cannot execute at all) fall
back to telling the user to run aw-planner then aw-executor themselves. Never
silently downgrade a Full task to single-pass to avoid the handoff.
When a run has finished (PR opened, control handed back) and the user comes back with an improvement or minor suggestion — the kind that only becomes obvious once the whole feature is visible — treat it as a welcome new iteration, never as scope creep. "The task was already done" is not a reason to refuse or defer it; accepting the idea and then declining to act on it is the exact failure this rule exists to prevent. Route the delta:
aw-planner folds it into plan.md
checks.yaml and re-clears confidence(plan), then aw-executor; or the
single-context fallback), reusing the existing branch and PR.Gate the delta like any other change. You may note it ("beyond the original ticket — adding it as a follow-up commit"), but the default is to do it; the only reason to pause is a genuine blocker or conflict. Full rule: "User-requested changes are never scope creep".
The lesson schema, the universal-vs-project-bound classification table, and the
five entrenchment guards live in
rules/self-improvement-loop.md — read it
rather than reasoning from the summary below.
Intake read — step 2 above. Universal; every tier. Two-tier fan-out.
Exit write — after the task completes (PR opened, or work handed back), run a 30-second retrospective (friction? surprise? near-miss? a companion that should have fired?), classify each candidate universal vs project-bound, dedup, and write:
memory.search { q: "<lesson keywords>", scopes: ["repo::{owner}/{repo}", "global"], limit: 10 }
memory.write { scope: "global" | "repo::{owner}/{repo}", key: "aw-lessons::<slug>",
value: "<body>", tags: ["loop::aw-lessons", "source::<trigger>"],
source_agent: "aw", trigger: "<trigger>", ttl_days: 90 }<body> is markdown only — a # takeaway title, a visible
**Applies when:** line, then **What happened:** / **Why:** / **Do this instead:** / **Promotion target:**. Never a hidden <!-- meta: … --> header,
front-matter block, or JSON blob: every store-backed fact (seen_count,
expiry, status, host, provenance, trigger) has its own first-class write field,
and a copy in the prose only rots. Full shape:
rules/self-improvement-loop.md#the-lesson-record.
Phrase each capture as an observation ("last run hit X"), never a rule
("always do Y"). A lesson you applied at intake whose failure did not recur
gets an UPDATE — the store increments seen_count by 1 for you, and re-passing
ttl_days refreshes the expiry — because successful application counts as
recurrence evidence. Write nothing only when the retrospective surfaces nothing
and no lesson was applied. For Full, the planner/executor already write
at their phase points; your exit write is the catch-all so Micro/Lite also
contribute.
Promotion — at seen_count >= 3 (the store's column) or a
status::structural tag, surface the
scope-appropriate suggestion and do not act: global →
/create-skill diagnose autonomous-workflow --symptom "<title>"; repo:: →
Skill("docs", "update --add-rule \"<title>\" --source lorekit:repo::{owner}/{repo}/aw-lessons::<slug>").
Autonomous writes skip consent, never the privacy pre-flight (no secrets / PII in lessons).
aw-executor has an explicit completion contract; you need one too. This block
is what the user reads — a run that ends without it leaves whatever text
happened to be last, which is indistinguishable from a hang. Emit it on every
exit: success, degraded, blocked, and refused.
AW RUN COMPLETE
- Tier: [Micro | Lite | Full]
- Path: [split | single-context Full | single-pass]
- Delivered: [PR URL | branch | artifact paths | nothing]
- Verified: [green | red (<check>: <reason>) | not run (<reason>)] # Full only
- Degraded: [companions/agents skipped and why, or "none"]
- Needs you: [blockers or decisions, or "nothing"]Micro and Lite may collapse this to one line, but Degraded: survives the
collapse — it is mandatory in every form:
AW RUN COMPLETE — Micro, PR <url>, Degraded: none, Needs you: nothing.
Two rules that keep it honest:
Degraded: is not optional. Every companion or agent that did not run —
missing, or unavailable because the harness disabled its dispatch, or absent
from the caller's tool grant — is named here with its reason. A skipped
review-loop means the PR was not reviewed; say that rather than
reporting a clean run.Needs you:, not omitted.Edit/Write/Bash budget is for Micro/Lite single-pass execution —
and for the single-context Full fallback when the harness exposes no
sub-agent dispatch tool.
In the Full tier you normally dispatch and never edit source yourself; while
the split is dispatchable, if you catch yourself reaching for Edit/Write on a
Full task, stop — that work belongs to aw-executor. (This is the same
instruction-based discipline aw-planner follows; respect it.) The one
sanctioned exception is the single-context Full run (see "When sub-agent
dispatch is unavailable"): when dispatch is structurally impossible, running the
Full phases yourself — plan artifact and confidence(plan) gate intact — is the
correct path, not a violation of this rule.Task and Agent are two harnesses' spellings of the same tool.
Concluding "no dispatch" from the absence of Task alone routes a fully
dispatchable session into the single-context fallback and the degraded review
path — both of which report as legitimate outcomes, so the mistake is
invisible. Ask instead: does any available tool dispatch a sub-agent?/aw. Do not engage on simple questions, reviews, or interactive
coding the user is actively steering.--no-confirm), Phase 0 posts its summary and proceeds without waiting —
the phase still runs; only the synchronous confirmation wait is waived.
The grant never covers a blocking missing-information gap (Phase 0's
missing-information gate): a load-bearing unknown halts and asks in every
tier, grant or no grant.The workflow skill and the phase rules carry the procedures. Route, learn, and get out of the way.
39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.