Content
50%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The skill has a coherent three-phase wizard structure with a useful size-limit table and an explicit question bank, but it undercuts itself: the document ends mid-code-fence with no validation phase, and it violates its own progressive-disclosure rules by inlining a ~40-line example while shipping no bundle files. Guidance for Phases 2–3 relies on placeholders and topic labels rather than executable instructions.
Suggestions
Complete the truncated Phase 3 example (close the code fence) and add a Phase 4 validation step that checks generated artifacts against the size-limit table, with a fix-and-recheck loop.
Move the embedded testing-skill example to a bundled template file (e.g., templates/testing-skill.template.md) and reference it with "See:", following the skill's own "Never embed examples >20 lines" rule.
Replace placeholders like "[Language-specific test template]" and topic labels like "Language-specific best practices / Common gotchas" with concrete instructions, examples, or references to bundled per-language template files.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly dense and functional (question banks, file paths, a size-limit table), but includes redundancy — the 100/500/400-line limits appear in both the table and the directory-structure comments — and re-explains Anthropic's progressive-disclosure tier pattern ("Loaded at startup for ALL installed skills", "Core workflow — NOT reference material") that is background rather than instruction. Not a 2 because there is no concept padding (no 'what a skill is' explanations); not a 4 because the tier diagram and duplicated limits could be trimmed. | 3 / 5 |
Actionability | Concrete elements exist — exact output paths (".github/copilot-instructions.md", ".github/instructions/[language].instructions.md") and a full question bank — but the generation phases lean on placeholders ("[Language-specific test template]", "[Stack-specific examples]", "[PROJECT_STACK]") and describe rather than instruct ("Language-specific best practices", "Common gotchas"). Not a 4 because the template placeholders are pseudocode rather than executable guidance; above 2 because file paths and question wording are directly usable. | 3 / 5 |
Workflow Clarity | Phases 1→2→3 are clearly sequenced, and the size-limit table's "Action if Exceeded" column gestures at enforcement, but there is no explicit post-generation validation step, and the workflow ends abruptly mid-example inside an unclosed code fence ("[Stack-specific examples]" is the last line) — Phases 4+ (verification, feedback loop) are simply absent. Meets anchor 3: sequence present but checkpoints missing; not a 2 because the sequencing that exists is coherent, and the truncation plus missing validation block any higher score. | 3 / 5 |
Progressive Disclosure | Sections are well organized (Purpose, Size Limits, When to Use, Workflow phases), but the ~40-line embedded testing-skill example is exactly the content the skill's own rule says to extract ("Never embed examples >20 lines in SKILL.md — use `See: examples.md`"), and the skill ships no references/, templates/, or scripts/ files at all — the "Always create templates/ subdirectory" rule is stated but not practiced. Meets anchor 3: some structure, content that should be separate is inline; not a 4 because the largest block of content violates the skill's own disclosure rules. | 3 / 5 |
Total | 12 / 20 Passed |