CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-development

Use when implementing any feature or bugfix, before writing implementation code

54

1.00x
Quality

55%

Does it follow best practices?

Impact

100%

1.00x

Average score across 1 eval scenario

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./eval/local/skills/benchmarks/dependency/superpowers/test-driven-development/SKILL.md

The canonical home for this skill is test-driven-development in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an unusually actionable TDD guide: concrete Good/Bad code, exact commands, worked examples, and a verification loop with explicit error-recovery guidance. Its weaknesses are redundancy — the rationalization rebuttals appear three times in different forms — and a dangling reference to a testing-anti-patterns.md file that is not in the bundle, which undermines the progressive-disclosure structure of an otherwise well-sectioned document.

Suggestions

Fix the broken reference: either add testing-anti-patterns.md to a references/ directory or remove the 'Testing Anti-Patterns' section's link, since the file does not exist in the bundle.

Consolidate the triple-repeated rationalization content ('Why Order Matters', 'Common Rationalizations', 'Red Flags') into one section — the same excuses are rebutted three times, costing tokens without adding guidance.

Consider moving the rationalization table and anti-patterns detail into the referenced file so SKILL.md stays a lean overview of the cycle, per the red flags list already being a summary.

DimensionReasoningScore

Conciseness

The TDD cycle, code examples, and checklist are tight and actionable, but the same anti-rationalization content is repeated three times ('Why Order Matters' prose, the 'Common Rationalizations' table covering the identical excuses — 'I'll test after', 'deleting X hours is wasteful', 'tests after achieve the same goals', 'already manually tested' — and the 'Red Flags' list repeating them again), plus a token-heavy graphviz diagram that a four-item list would convey. Mostly efficient with clear tightening opportunities, so not anchor 2's 'several padded sections' nor anchor 4's 'minor instances'.

3 / 5

Actionability

Copy-paste-ready TypeScript examples with explicit Good/Bad pairs for both the test and the implementation, exact commands ('npm test path/to/test.test.ts'), a fully worked bug-fix example showing expected FAIL/PASS output, a verification checklist, and a troubleshooting table. This covers the common cases the way the anchor-5 example does.

5 / 5

Workflow Clarity

The red-green-refactor sequence is explicit with mandatory verification checkpoints at each phase, failure-mode feedback loops ('Test passes? You're testing existing behavior. Fix test.', 'Test errors? Fix error, re-run until it fails correctly', 'Test fails? Fix code, not test.'), a pre-completion checklist, and 'Confirm' criteria distinguishing expected failure from typos — matching the anchor-5 pattern of clear sequence plus validation plus error-recovery loops.

5 / 5

Progressive Disclosure

The body is a single ~370-line monolith with clear section headers but no working bundle: the one reference — 'read [testing-anti-patterns.md](testing-anti-patterns.md)' — points to a file that does not exist (no references/ directory or sibling file is present), and large reusable content (the rationalization tables, anti-patterns) is inlined rather than split out. That sits at anchor 3 ('some structure, references present but problematic, content that should be separate is inline') — above anchor 2 because sections are well-organized and the reference is clearly signaled, below 4 because the sole reference is broken.

3 / 5

Total

16

/

20

Passed

Description

32%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a pure trigger clause: it tells Claude when to load the skill but not what the skill does, and it triggers so broadly that it overlaps nearly all coding tasks. It also omits the skill's own vocabulary (test, TDD, test-first), missing the terms a user would most naturally say. Adding a one-line 'what' (e.g., 'Write a failing test first, then minimal implementation code, following red-green-refactor') and narrowing the trigger would fix most of this.

Suggestions

State the 'what': lead with concrete actions, e.g. 'Enforce test-driven development: write a failing test first, implement minimal code to pass it, then refactor (red-green-refactor).'

Include the skill's natural trigger vocabulary — 'test', 'tests', 'TDD', 'test-first', 'write the test first' — so users who ask in those terms surface this skill.

Narrow or differentiate the trigger so it does not compete with every general coding request; the current 'any feature or bugfix' clause conflicts with virtually all implementation work.

DimensionReasoningScore

Specificity

The description names the domain ("implementing any feature or bugfix") but describes zero actions — it never says what the skill actually does (write failing tests first, red-green-refactor). It falls between anchor 1 (pure abstract language, which this is not — the wording is concrete) and anchor 3 (which requires 1-2 concrete actions, absent here).

2 / 5

Completeness

This is only a 'when' clause with no 'what' — the exact anchor-2 case ('only when is present without what'). It explicitly says when to use the skill but never states what the skill does, so it cannot reach 3 (which requires a clear 'what') and is above 1 because the 'when' is explicit and specific.

2 / 5

Trigger Term Quality

"implementing any feature or bugfix" and "before writing implementation code" are natural phrases users say, but the obvious trigger vocabulary for this skill — "test", "tests", "TDD", "test-first", "write tests" — is entirely missing, matching the 'some relevant keywords but missing common variations or synonyms' anchor.

3 / 5

Distinctiveness Conflict Risk

"Use when implementing any feature or bugfix" would fire on essentially all coding work, creating high overlap risk with any implementation, refactoring, or code-review skill. It is not anchor-1 generic ('helps with code and documents') because it is anchored to a real coding activity, but it is very broad.

2 / 5

Total

9

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
rpamis/comet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.