CtrlK
BlogDocsLog inGet started
Tessl Logo

subagent-driven-development

Use when executing implementation plans with independent tasks in the current session

48

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./eval/local/skills/benchmarks/dependency/superpowers/subagent-driven-development/SKILL.md

The canonical home for this skill is subagent-driven-development in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an unusually strong orchestration guide: fully sequenced workflow with real validation gates and feedback loops, concrete commands, status-handling taxonomy, and hard-won failure lessons rendered as specific rules. Its weaknesses are moderate redundancy (duplicated graph edges, an Advantages section that restates earlier material, a long example transcript) and a progressive-disclosure failure — the two dispatch prompt templates it repeatedly points to are absent from the bundle, so the process cannot be executed as written without them.

Suggestions

Ship the referenced bundle files implementer-prompt.md and task-reviewer-prompt.md (or inline their essential contracts), since the process graph, Prompt Templates section, and File Handoffs section all depend on them.

Cut redundancy: render each graphviz digraph once (edges only, drop the duplicated node/edge listing), trim the Advantages section to the few points not already covered by the core principle and Red Flags, and compress the Example Workflow to one task cycle plus the fix loop.

Make model selection actionable by naming concrete model tiers/ids for the implementer, reviewer, and final-review dispatches instead of relative labels like "a fast, cheap model".

DimensionReasoningScore

Conciseness

The body is dense and operational — it teaches non-obvious process economics ("a real session's dispatch hit 42k chars", "Turn count beats token price") rather than concepts Claude already knows — but several sections are padded: the two graphviz digraphs list every node and then repeat every edge as text, the Advantages section restates the core principle and Red Flags in softer form, and the Example Workflow transcript runs ~60 lines to illustrate a loop the Process section already specifies. This is anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened'), above anchor 2 because the fat is redundancy, not explanation of known concepts.

3 / 5

Actionability

Guidance is largely executable: exact commands ("scripts/review-package BASE HEAD", "scripts/task-brief PLAN_FILE N", the ledger cat path), a literal ledger line format, a five-part dispatch composition recipe, a four-status handling taxonomy, and a three-element fix-report checklist — and the referenced scripts exist in the bundle with matching interfaces. It falls short of anchor 5 because the model-selection tiers name no concrete models ("a fast, cheap model" is not actionable without ids), and the two dispatch prompt templates the process depends on (implementer-prompt.md, task-reviewer-prompt.md) are referenced but absent from the bundle, so the core dispatch artifact cannot actually be used as written.

4 / 5

Workflow Clarity

The multi-step process is fully sequenced (read plan → pre-flight conflict scan → per-task dispatch/question/review/fix loop → final whole-branch review) with explicit validation checkpoints at every stage: review gates with two required verdicts, re-review loops until approved, ⚠️-item resolution before task completion, BLOCKED escalation rules, and a fix contract requiring test evidence before re-review. This matches anchor 5 ('clear sequence with explicit validation steps; feedback loops for error recovery; checklists') — the Red Flags list is a genuine checklist and the ledger is an explicit recovery mechanism after compaction.

5 / 5

Progressive Disclosure

Section structure is good and the scripts are properly externalized (review-package, task-brief, sdd-workspace all exist and are invoked by name), but the two central references — "./implementer-prompt.md" and "./task-reviewer-prompt.md", linked from the Prompt Templates section and the process graph — are missing from the bundle, leaving dangling links to the skill's most important artifacts; the final-review reference (../requesting-code-review/code-reviewer.md) is likewise external and unverifiable. This is anchor 3 ('references present but not clearly signaled / problematic'): the split is conceptually right, but navigation to the key materials is broken, which blocks anchor 4's 'references mostly clear'.

3 / 5

Total

15

/

20

Passed

Description

36%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a pure trigger clause: it answers 'when' clearly but omits 'what' entirely, forcing the model to infer the skill's mechanism from the name alone. Trigger phrasing is serviceable but thin on natural synonyms and leans on internal jargon ('current session'). Adding one sentence of concrete capability (per-task implementer subagents, spec+quality review gates, final whole-branch review) would lift specificity and completeness substantially.

Suggestions

Add a 'what' statement before the trigger clause, e.g. "Execute implementation plans by dispatching a fresh implementer subagent per task with spec-compliance and code-quality review gates and a final whole-branch review."

Broaden trigger coverage with natural user phrasings like "execute this plan", "run the plan tasks", or "implement the plan with subagents" so it fires on how users actually ask.

Replace or gloss the jargon "in the current session" with a user-facing distinction (e.g. "same-session, subagent-driven execution") to reduce overlap with a parallel-session plan-execution skill.

DimensionReasoningScore

Specificity

The description names its domain ("executing implementation plans with independent tasks in the current session") but states no concrete actions — there is no verb describing what the skill actually does (dispatch subagents, review, fix). It sits above anchor 1 ('entirely vague') because the domain is real and scoped, but matches anchor 2 ('names the domain but actions are minimal or generic') since the operational capability is entirely unstated.

2 / 5

Completeness

The 'when' half is explicitly answered ("Use when executing implementation plans with independent tasks in the current session") but the 'what' is absent — the description never states that the skill dispatches per-task implementer/reviewer subagents with review gates. This matches anchor 2's second disjunct exactly ('only when is present without what', e.g. "Use when working with documents"); it is not 3 because the what-component is not merely weak — it is unstated, and not 1 because the when half is concrete and complete.

2 / 5

Trigger Term Quality

"executing implementation plans" and "independent tasks" are phrases a user might plausibly say, but common natural variations are missing ("execute this plan", "run the plan", "work through the tasks", "subagent", "delegate tasks"), and "in the current session" is internal jargon aimed at disambiguating a sibling skill rather than a user utterance. This matches anchor 3 ('some relevant keywords but missing common variations or synonyms') and is below anchor 4's 'good keyword coverage'.

3 / 5

Distinctiveness Conflict Risk

"executing implementation plans" overlaps heavily with generic plan-execution and the sibling executing-plans skill; the qualifiers "independent tasks" and "in the current session" do narrow it, but the subagent-driven niche that actually distinguishes this skill is never mentioned. This is anchor 3 ('somewhat specific but could still overlap with similar skills') — not 4, since the disambiguating feature (same-session subagent dispatch vs. parallel-session execution) is only half-captured.

3 / 5

Total

10

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing, 1 suspicious

Warning

Total

15

/

16

Passed

Repository
rpamis/comet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.