CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-dev

Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/test-driven-dev/SKILL.md

The canonical home for this skill is tdd in mattpocock/skills

SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A clearly structured, opinionated TDD workflow with strong sequencing, checklists, and concrete rules of engagement. Its weaknesses are a padded Philosophy section that re-explains what Claude already knows, missing error-recovery branches, and five dead file references that promise content the bundle does not deliver.

Suggestions

Trim the Philosophy section to 3-4 directive bullets (e.g. 'Test behavior via public interfaces only; if a refactor breaks a test, the test was testing implementation') and cut the discursive prose in the Anti-Pattern section.

Add the referenced files (tests.md, mocking.md, deep-modules.md, interface-design.md, refactoring.md) to the bundle — or remove the dead links and inline the one or two that carry essential guidance.

Add brief error-recovery branches to the loop: what to do when a new test passes immediately (delete or question the test), when the plan proves wrong mid-cycle, and when growing duplication signals it is time to refactor mid-loop.

DimensionReasoningScore

Conciseness

The Philosophy section spends three paragraphs explaining concepts Claude already knows (good vs. bad tests, mocking internal collaborators, tests breaking on refactor), and the Anti-Pattern section is similarly discursive ('You outrun your headlights'). This is 'mostly efficient but includes some unnecessary explanation' rather than anchor 4 — the Philosophy section alone could be trimmed to a few directive bullets without losing guidance value.

3 / 5

Actionability

As an instruction-only skill the guidance is concrete and directive: 'Write ONE test that confirms ONE thing', 'Only enough code to pass current test', 'Never refactor while RED', a specific question to ask the user ('What should the public interface look like?'), and a worked behavior-naming example ('user can checkout with valid cart'). Not score 5 because no example test or RED/GREEN cycle in a real framework is shown and the referenced example files (tests.md) do not exist, leaving minor gaps.

4 / 5

Workflow Clarity

The four-step workflow (Planning → Tracer Bullet → Incremental Loop → Refactor) is clearly sequenced with three checklists and explicit cycle validation ('test fails' → 'test passes', 'Run tests after each refactor step'). Not score 5 because there are no error-recovery branches — e.g. what to do when a new test unexpectedly passes, when implementation reveals the plan was wrong, or when to abandon a cycle — so feedback loops exist but only on the happy path.

4 / 5

Progressive Disclosure

The body is well-sectioned and its inline references (tests.md, mocking.md, deep-modules.md, interface-design.md, refactoring.md) are clearly signaled and one level deep, which is better organized than anchor 2. However, none of the five referenced files exist in the bundle (no references/, scripts/, or assets/ directories), so navigation leads nowhere and the promised detail content is simply missing — a real organizational defect that keeps it below anchor 4.

3 / 5

Total

14

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-formed description with an explicit 'Use when' clause and rich, natural trigger vocabulary that clearly matches the skill's TDD niche. The main weakness is that the 'what' stays at the process level rather than naming concrete capabilities.

Suggestions

Sharpen the 'what' with concrete deliverables, e.g. 'Builds features and fixes bugs via the red-green-refactor loop: write one failing behavior test, minimal implementation to pass, then refactor.'

Add one or two common synonyms such as 'unit tests' or 'write the tests first' to the trigger list to broaden natural activation phrasing.

DimensionReasoningScore

Specificity

The 'what' is stated at a process level ('Test-driven development with red-green-refactor loop') with implied actions ('build features or fix bugs'), which is comparable to 'names domain and 1-2 concrete actions, but not comprehensive'. Not score 4 because it never enumerates the concrete operations the way 'extracts text, fills forms, converts pages' does; not score 2 because 'red-green-refactor loop' is more concrete than a bare domain name like 'Processes PDF files'.

3 / 5

Completeness

Explicitly answers both: 'what' — 'Test-driven development with red-green-refactor loop', and 'when' — a full 'Use when...' clause with multiple concrete trigger phrases (build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, asks for test-first development). Matches the anchor-5 example structure; anchor 4's 'when could be more explicit' does not apply since the when-clause lists four specific triggers.

5 / 5

Trigger Term Quality

Strong natural keyword coverage: 'TDD', 'red-green-refactor', 'integration tests', 'test-first development', 'build features', 'fix bugs' — phrases a user would plausibly say. Not score 5 because common variations like 'unit tests', 'write the tests first', or 'test coverage' are absent, so coverage is good but not comprehensive.

4 / 5

Distinctiveness Conflict Risk

TDD/red-green-refactor is a clear niche with distinct triggers, so conflict risk with unrelated skills is low. Not score 5 because 'wants integration tests' and the broad 'build features or fix bugs' phrasing create minor overlap with general testing and coding skills despite the 'using TDD' qualifier.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 missing

Warning

Total

15

/

16

Passed

Repository
zebbern/claude-code-guide
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.