CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-tdd

Build a behavior change with observed red, minimal green, and measured test consolidation

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-tdd/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, disciplined instruction body that encodes TDD as enforceable gates: an explicit hard gate, a numbered rule with red/green checkpoints, a quantified consolidation protocol with a required ledger, and a rotation heuristic for stuck cycles. Its only weaknesses are implicit output formats (no ledger example) and references to files that do not exist in this skill's bundle.

DimensionReasoningScore

Conciseness

The body is lean throughout: it assumes Claude already knows TDD and adds only policy — 'If a test passes before implementation, it is not red evidence', 'Review a slowdown only when it exceeds both 20 percent and 100 ms'. There is no padding, no conceptual explanation of what tests or refactoring are, and every sentence constrains behavior.

5 / 5

Actionability

Guidance is highly concrete for an instruction-only skill: a numbered six-step rule, named ledger fields ('old_test', 'behavior', 'replacement', 'mutant', 'red_observed', 'baseline_ms', 'candidate_ms', 'reason'), and exact measurement protocol (one warm-up, five isolated runs, median). It stops short of a 5 because no example of the ledger or a red-evidence record is given, so the exact output format is left implicit.

4 / 5

Workflow Clarity

The sequence is explicit with validation checkpoints at each gate: run the focused test and record the expected failure and exit status, implement only after red, run the focused test then the affected suite, refactor only while green. Feedback loops are present — repairing a test that errors 'until it fails on the missing behavior', and stopping after two failed attempts to re-check the boundary — satisfying the rubric's error-recovery requirement.

5 / 5

Progressive Disclosure

The skill is a short, single file with well-organized sections (rule, consolidation, strategy rotation) and no bundle directory, so there is nothing that should be split out. It loses a point because the body references files that are not part of this skill's bundle — 'skills/blocks/engineering-method-selection.md' and 'THIRD_PARTY_NOTICES.md' — leaving two dangling paths a reader cannot resolve locally.

4 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A terse, distinctive description that names the core TDD actions concretely, but it never says when to use the skill and leans on cycle jargon ('red', 'green') instead of the natural terms users actually say. The 'when' guidance lives in a separate frontmatter 'trigger:' field rather than in the description itself, which caps its completeness.

Suggestions

Fold the frontmatter trigger condition into the description so it answers 'when' on its own, e.g. append 'Use when a feature, bug fix, or behavior change requires test-first evidence; do not use for documentation-only changes.'

Add natural trigger terms and synonyms users would say, such as 'test-driven development', 'TDD', 'write the test first', or 'unit tests', to improve keyword coverage.

Spell out the cycle in plainer language alongside the jargon (e.g., 'write a failing test first, implement the smallest change to pass it, refactor with tests green') so the what is concrete to non-TDD-fluent requesters too.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'Build a behavior change', 'observed red, minimal green', 'measured test consolidation' — each naming a distinct step of the TDD cycle with modifiers that constrain how it is done. It falls short of a 5 because coverage is compressed into jargon shorthand and omits parts of the cycle (e.g., refactor), leaving minor gaps.

4 / 5

Completeness

The 'what' is clearly stated — build a behavior change through observed red, minimal green, and measured test consolidation — but the description itself contains no 'Use when...' clause or equivalent explicit trigger guidance, capping completeness at 3 per the rubric. (The frontmatter has a separate 'trigger:' field, but the description field evaluated here does not answer 'when' on its own.)

3 / 5

Trigger Term Quality

It contains some relevant keywords ('behavior change', 'red', 'green', 'test consolidation') but misses the natural phrases most users say, such as 'test-driven development', 'TDD', 'write tests first', or 'unit tests'. It sits above 2 (more than one or two generic keywords) but below 4, which requires good coverage of natural terms users would actually say.

3 / 5

Distinctiveness Conflict Risk

The red/green/consolidation language carves out a clear TDD niche that is unlikely to trigger for unrelated skills, but the opening phrase 'Build a behavior change' is broad and the description lacks the explicit 'test'/'TDD' naming that would fully separate it from general coding or testing skills, keeping it just below a 5.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.