CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-development

Use when implementing any feature or bugfix, before writing implementation code

48

Quality

51%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/superpowers/skills/test-driven-development/SKILL.md

The canonical home for this skill is test-driven-development in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strongly actionable, well-sequenced TDD skill whose Red-Green-Refactor workflow with per-phase verification and recovery loops is exemplary. Its weaknesses are repetition of the same anti-rationalization message across three sections, a non-rendering graphviz diagram, and a broken link to a writing-good-tests.md file that is not in the bundle.

Suggestions

Create the referenced writing-good-tests.md (or remove the link) — the four bullets that follow the link are a natural seed for that file, which would also slim the body.

Merge the Red Flags list with the Common Rationalizations table, since nearly every red flag duplicates a row of the table, and drop or shrink the graphviz diagram to a one-line cycle description.

Replace fragmentary code examples (the `// ...` in submitForm, the `// YAGNI` stub) with complete, runnable versions to reach copy-paste-ready coverage.

DimensionReasoningScore

Conciseness

The body is mostly tight and imperative, but the same enforcement message is repeated across three sections ("Delete it. Start over" in The Iron Law, "Keep as reference, write tests first... Delete means delete" in Common Rationalizations, and "Keep as reference" or "adapt existing code" in Red Flags), and the ~20-line graphviz `dot` diagram renders as raw code in markdown, adding tokens without adding instruction. Not level 4 because the duplication and the decorative diagram are trimming candidates beyond 'minor'; not level 2 because every section still carries operational content rather than padding.

3 / 5

Actionability

Concrete, executable guidance dominates: real commands (`npm test path/to/test.test.ts`), complete TypeScript test/implementation examples with Good/Bad contrasts, a worked bug-fix walkthrough, a verification checklist, and a troubleshooting table. Not level 5 because some examples are fragments — the GREEN example for retry is fully executable, but `submitForm` ends in `// ...` and the Bad GREEN example is a stub with `// YAGNI` — and the link to writing-good-tests.md leads nowhere.

4 / 5

Workflow Clarity

The Red-Green-Refactor cycle is explicitly sequenced with a mandatory verification command after each phase ("Verify RED - Watch It Fail... MANDATORY. Never skip", "Verify GREEN... MANDATORY"), includes feedback loops for error recovery ("Test errors? Fix error, re-run until it fails correctly", "Test fails? Fix code, not test"), and closes with a per-item verification checklist plus a "When Stuck" recovery table. This matches the level-5 anchor: clear sequence, explicit validation steps, and error-recovery loops.

5 / 5

Progressive Disclosure

The single external reference is well-signaled inline ("read [writing-good-tests.md](writing-good-tests.md) for the rules that keep tests honest") and one level deep, but the file does not exist — no references/ directory or writing-good-tests.md is present anywhere in the skill bundle, so the link is broken. The body is also ~315 lines with content (the Good Tests rules and the Common Rationalizations table) that reads like material meant for that missing file. Not level 4 because a dangling reference defeats navigation; not level 2 because section structure is clear and there is no monolithic wall of text.

3 / 5

Total

15

/

20

Passed

Description

32%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has an explicit and well-timed 'Use when' trigger, but omits the 'what' entirely — it never mentions tests, TDD, or red-green-refactor — making it indistinguishable from a generic coding assistant and prone to firing on every implementation task. Trigger vocabulary is limited to 'feature' and 'bugfix' with none of the test-related terms users would naturally use.

Suggestions

Lead with what the skill does, e.g. 'Implements features and bugfixes using strict test-driven development: write the failing test first, watch it fail, write minimal code to pass.' before the 'Use when' clause.

Add natural trigger terms users would actually say — test, tests, TDD, test-first, unit tests, red-green-refactor — so the skill surfaces when the user mentions testing.

Narrow the trigger scope or state its exclusivity (e.g. 'Use whenever production code is being written or changed, superseding direct-implementation approaches') to reduce conflict with general implementation workflows.

DimensionReasoningScore

Specificity

The description names the domain ("implementing any feature or bugfix") but states no action at all — a reader cannot tell from it that the skill mandates test-first development. Matches the level-2 anchor 'Names the domain but actions are minimal or generic'; not level 3 because it lists zero concrete actions (no 'write failing tests first', 'verify red/green'), and not level 1 because the domain is named and the timing constraint is concrete.

2 / 5

Completeness

Only the 'when' is present ("Use when implementing any feature or bugfix, before writing implementation code") — the 'what' (what the skill actually does) is missing entirely, matching the level-2 anchor 'only when is present without what'. Not level 3 because that anchor requires a clear 'what'; not level 1 because the 'when' clause is explicit and specific about timing.

2 / 5

Trigger Term Quality

"implementing any feature or bugfix" uses natural phrases users say ("implement a feature", "fix a bug"), but the vocabulary users would most plausibly use for this skill — test, tests, testing, TDD, test-first, unit test — is entirely absent. Matches level 3 'Some relevant keywords but missing common variations or synonyms'; not level 4 because the missing terms are the skill's own subject matter, not just edge variations.

3 / 5

Distinctiveness Conflict Risk

"any feature or bugfix" covers essentially every coding task, so this description would trigger alongside virtually any implementation, refactoring, or testing skill — matching level 2 'Very broad; high overlap risk with many similar skills'. Not level 3 because the only distinguishing signal (write tests first) is absent; not level 1 because it is at least anchored to feature/bugfix work rather than being fully generic like 'Helps with code'.

2 / 5

Total

9

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.