Scaffolds, reviews, upgrades, or diagnoses agent skills against best-practice frontmatter, progressive disclosure, token-aware structure, and the agent-skills.git symlink + inventory wiring. Modes: `scaffold` (default — new skill), `review` (audit existing skill), `upgrade` (split a single-file skill into multi-file), `diagnose` (retrospective failure analysis that emits a confidence-gated unified diff against any skill declaring a diagnostic surface). Triggers on "create a skill", "scaffold a skill", "new SKILL.md", "review this skill", "audit my skill", "upgrade this skill", "split this skill", "diagnose this skill", "why did the skill miss this", "/create-skill".
72
91%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Author new agent skills — and audit existing ones — against the best practices distilled from Anthropic's official Skill authoring guide and the patterns already in this repo. Output is a complete skill directory plus the agent-skills.git symlink wiring and inventory updates.
This
SKILL.mdis a thin index. Detailed authoring rules live inrules/*.md, worked examples inreferences/*.md, literal scaffolding templates intemplates/*.md— all load on demand. Reading them all up-front would burn tokens you do not need yet.
Parse $ARGUMENTS (first token) and detect the mode:
| Mode | Default | Trigger |
|---|---|---|
scaffold | yes | Default. "create", "scaffold", "new skill", or no mode argument. |
review | "review", "audit", "check this skill", or $0 == "review". | |
upgrade | "upgrade", "split", "convert to multi-file", or $0 == "upgrade". | |
diagnose | "diagnose", "why did miss this", or $0 == "diagnose". |
If the user typed a path or skill name as $ARGUMENTS, treat it as the
target for review/upgrade/diagnose; for scaffold it is the proposed
skill name.
State the detected mode and target in one line before continuing. Example:
Mode: scaffold
Target: skills/<category>/<proposed-name>/A seven-phase pipeline. Each phase has a gate; do not proceed until it passes.
| Phase | Name | Gate |
|---|---|---|
| 0 | Requirements | User confirmed name, description, modes, structure choice, target runtime |
| 1 | Structure decision | Single-file vs multi-file decided with reasoning |
| 2 | Frontmatter draft | Name + description + flags pass validation |
| 3 | File generation | All planned files written, none over budget |
| 4 | Wiring & inventories | Symlinks created (if local-dev), CLAUDE.md + README.md updated |
| 5 | Self-check | Mechanical pre-pass (scripts/validate-skill.mjs) is PASS, and every judgment item in rules/quality-checklist.md passes |
| 6 | Evaluation | ≥ 3 eval prompts written and at least one with-skill run observed, or the user explicitly waived it |
Ask the user — in one message, batched, so they answer once:
name: field be?disable-model-invocation: true),
model-invokable (default), or hidden background (user-invocable: false)?
See rules/invocation-control.md.$ARGUMENTS? Positional ($0, $1)?
None?allowed-tools pre-approve any specific tools (e.g.
Bash(git *))? Default: leave unset.claude-code (default, full frontmatter) or
portable (Skills API / claude.ai upload, six spec fields only)? See
rules/frontmatter.md § Portability profile.evals/ now rather
than after the fact — see rules/evaluation.md.Confirm the answers back to the user verbatim before moving on. Do not guess any of these.
Decide: single-file or multi-file? Apply this decision flow before
generating anything — see rules/structure-decision.md for the full rubric.
Quick decision table:
| Signal | Pick |
|---|---|
| Body fits comfortably under 200 lines | Single-file |
| Body would exceed 500 lines (the hard cap) | Multi-file |
| 3+ orthogonal concerns (e.g. naming + architecture + tests) | Multi-file (one rule per concern) |
| Worked examples > 100 lines | Move to references/ |
| Reusable boilerplate the skill emits literally | Move to templates/ |
| One mode and one concern | Single-file |
Output the chosen layout as a tree before writing files.
In this repo the directory is nested one level under a category (workflow/, quality/, delivery/, testing/, design/, analysis/, or authoring/ — see rules/repository-conventions.md):
skills/<category>/<name>/
├── SKILL.md
├── rules/...
├── references/...
└── templates/...Draft the YAML frontmatter using rules/frontmatter.md and
rules/description-writing.md. Validate before writing:
name is kebab-case, ≤ 64 chars, no reserved words (anthropic, claude).description is third-person, ≤ 1024 chars, includes both what and
when (trigger phrases), front-loaded with the most important keywords.disable-model-invocation: true is set if the user picked slash-only.argument-hint is set (unless the skill is user-invocable: false),
derived from the Modes (Q5) and Inputs (Q6) answers from Phase 0.
Mirror the actual modes / flags; use […] for optional, <…> for
placeholders, | for alternatives. If the skill takes no arguments,
emit argument-hint: '' explicitly rather than omitting the field.metadata.tags are populated (5–10 specific terms).Write each planned file. For each one:
SKILL.md — start from templates/SKILL.minimal.md (single-file),
templates/SKILL.multi-file.md (index pattern), or
templates/SKILL.portable.md (target runtime is portable). Keep body
≤ 500 lines.rules/<concern>.md — start from templates/rule.md, one file per
concern, loadable in isolation. TOC past 150 lines. Degrees-of-freedom
and checklist shapes: rules/workflow-patterns.md.references/<topic>.md — start from templates/reference.md. TOC
past 100 lines (Claude partial-reads long files; the TOC is the safety
net).templates/<artefact>.md — literal text the skill emits. No prose
meta-commentary inside templates.scripts/<name>.mjs — only for a deterministic check or transform
of the skill's own. Zero dependencies, ${CLAUDE_SKILL_DIR}-anchored,
a --self-test mode. See rules/scripts-and-assets.md.evals/evals.json, evals/triggers.jsonl — this skill's own test
prompts and trigger set, from templates/evals.json and
templates/triggers.jsonl. See rules/evaluation.md.After each file, verify:
If the user runs the local-dev symlink chain (the default for this repo),
follow rules/repository-conventions.md to:
skills/<category>/<name>/ (categories: workflow,
quality, delivery, testing, design, analysis, authoring).bash scripts/sync-symlinks.sh from the repo root to wire the
two-tier chain (~/.claude/skills/<name> → ~/.agents/skills/<name> →
<repo>/skills/<category>/<name>) — never ln -s by hand, and never
invoke the script with sh.readlink.CLAUDE.md (under the matching
### \/`` subsection, with the correct type marker).README.md and add the skill to the
"Repository Structure" tree at the bottom of the README.If the user is publishing the skill via npx skills add only, skip steps
2–3 but still update the inventories.
Run node ${CLAUDE_SKILL_DIR}/scripts/validate-skill.mjs <dir> [--portable]
first, then work through the remaining (judgment) items in
rules/quality-checklist.md. Treat any unchecked item as a defect — fix it
before declaring the skill done. Report the validator's own Self-check: line
verbatim — it prints its own pass/total, so never restate that count from
memory or invent one — then confirm the remaining (judgment) items passed.
On failure:
FAIL FM05 SKILL.md:3 — description must be non-empty and <= 1024 chars, got 1180
Self-check: FAIL — 1 failingFollow rules/evaluation.md: write and run at least 3 realistic test
prompts, observe at least one with-skill run, and check the repo eval
obligation table. Report Evaluation: PASS — 3 prompts run, with-skill navigation observed, or state explicitly that the user waived this phase.
For review mode, do not write any files. Read the target skill (the path
or skill name from $ARGUMENTS) and run the mechanical pre-pass, then work
through the judgment items:
node ${CLAUDE_SKILL_DIR}/scripts/validate-skill.mjs <dir> [--portable]; record every FAIL/WARN as evidence, then load rules/quality-checklist.md for the remaining (judgment) items.SKILL.md. If it has rules/, references/,
templates/, list each file with line count.Do not mutate the skill in review mode.
For upgrade mode, take a single-file skill and split it into multi-file:
SKILL.md.rules/, references/, templates/) and show
it to the user for approval before writing.SKILL.md with a one-line pointer + link to the new file.For diagnose mode, do not scaffold or review.
Analyse a session in which another skill executed and produced an
unsatisfactory result, identify which of that skill's gates should have
caught it, and emit a confidence-gated unified diff that hardens the target
skill against the same failure class.
The full procedure (seven steps, including the mandatory
confidence(analysis) ≥ 90 % gate before --apply), the report format,
and the hard rules live in rules/diagnose-mode.md.
Invocation:
/create-skill diagnose <target-skill-name> [--symptom "..."] [--scope <phase|companion>] [--apply] [--pr] [--no-write]The target declares its own diagnostic surface in
skills/<target>/rules/diagnostic-surface.md (skills) or
agents/<target>/rules/diagnostic-surface.md (agents) — phase model,
failure taxonomy, existing-guards table, source root, hard invariants.
Step 1 of Diagnose Mode disambiguates by checking both locations.
The contract spec is in rules/diagnostic-surface.md;
the scaffolding template a target drops into its own rules/ is
templates/diagnostic-surface.template.md.
If the target has not declared a surface, Diagnose Mode falls back to
inferring phases from the target body's H2 sections (SKILL.md for skills,
agents/<name>.md for agents) and warns the user once that fidelity is reduced.
Diagnose Mode never modifies user product code. It only proposes changes to the target's own source.
Self-improving skills. An orchestrator skill can close the loop further with
a two-tier self-improvement loop: a fast episodic-lessons tier
(LoreKit memory.* tools) feeding the slow diagnose tier via a recurrence gate.
The reusable recipe — including when NOT to add one — is in
rules/self-improvement-loop-pattern.md.
When a target declares a ## Lessons scope, Diagnose Mode reads it as evidence
(Step 2).
Load these on demand — do not preload them all.
| Phase | Files |
|---|---|
| 0 | rules/description-writing.md, rules/invocation-control.md |
| 1 | rules/structure-decision.md, rules/progressive-disclosure.md |
| 2 | rules/frontmatter.md, rules/description-writing.md, rules/arguments-and-injection.md |
| 3 | rules/token-economics.md, rules/anti-patterns.md, rules/arguments-and-injection.md, rules/scripts-and-assets.md, rules/workflow-patterns.md, plus templates in templates/ |
| 4 | rules/repository-conventions.md |
| 5 | rules/quality-checklist.md |
| 6 | rules/evaluation.md |
| diagnose | rules/diagnose-mode.md, rules/diagnostic-surface.md, plus the target's rules/diagnostic-surface.md |
| loop | rules/self-improvement-loop-pattern.md (adding a self-improvement loop to an orchestrator skill) |
| lens | rules/review-lens-contract.md + templates/lens.md (making an existing skill lens-eligible for pr-reviewer) |
references/skill-archetypes.md and references/good-vs-bad-examples.md
are optional — load only when the user asks for a worked shape or pair.
SKILL.md is a recurring token cost once loaded — write nothing Claude
already knows. See rules/token-economics.md.SKILL.md (loaded on trigger), supporting files
(loaded on demand). Keep references one level deep.rules/structure-decision.md.Skill() calls.rules/evaluation.md.rules/anti-patterns.md)description ("I can help you …").SKILL.md → a.md → b.md → c.md).anthropic, claude) in the name.disable-model-invocation, context, …) in a
skill targeting the portable six-field profile.A scaffold run is done when:
name and description validate against the rules in
rules/frontmatter.md.npx skills install path documented.CLAUDE.md and README.md added.node ${CLAUDE_SKILL_DIR}/scripts/validate-skill.mjs <dir> reports PASS.PASS.claude-code or portable) is recorded.A diagnose run is done when:
F-novel plus a
proposed new row).confidence(analysis) score recorded; --apply honored only at
≥ 90 % (final score, after Step 6.5's two-iteration refinement loop
if the initial score was below the gate)..agent/{branch}/diagnose-{target}.md (or
stdout with --no-write).--apply ran, user explicitly confirmed before git apply.39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.