CtrlK
BlogDocsLog inGet started
Tessl Logo

tdd

Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/tdd/SKILL.md

The canonical home for this skill is tdd in mattpocock/skills

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, opinionated body that gives concrete rules, a user-confirmation checkpoint, and crisp anti-pattern guidance without padding. Its one real defect is structural: it points readers to tests.md and mocking.md, which are absent from the skill bundle, so the promised detail is unreachable.

Suggestions

Add the referenced files (tests.md with good/bad test examples, mocking.md with mocking guidelines) to the skill bundle, or remove the links and inline the one or two examples that matter.

Add an explicit 'verify the test actually fails (red) before implementing' step to the Rules of the loop to close the validation gap in the loop sequence.

Trim concept re-explanations Claude already knows (the definition of a seam, the mechanics of tautological assertions) to tighten the body further.

DimensionReasoningScore

Conciseness

The body is lean and directive with no filler, no library tours, and no basic-concept padding; sections like 'Rules of the loop' are one line each. A few sentences re-explain concepts Claude already knows (e.g., defining 'seam', explaining what a tautological test is at length with examples like `expect(add(a, b)).toBe(a + b)`), which keeps it just below the 'every token earns its place' anchor.

4 / 5

Actionability

For an instruction-only skill the guidance is concrete: an exact question to ask ('What's the public interface, and which seams should we test?'), explicit rules ('Write the failing test first, then only enough code to pass it', 'One seam, one test, one minimal implementation per cycle'), and a named tool invocation ('call the Skill tool with "codebase-design"'). It misses anchor 5 because the concrete good-test and mocking examples are delegated to files that don't exist in the bundle, leaving minor gaps in executable specificity.

4 / 5

Workflow Clarity

The loop is clearly sequenced — agree seams with the user, red (failing test) before green (minimal implementation), one slice at a time — and includes an explicit checkpoint ('write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam'). It stops short of anchor 5: there is no explicit verify-the-test-fails step or error-recovery loop, and refactoring is delegated out without saying when review happens.

4 / 5

Progressive Disclosure

Sections are well organized and the references are clearly signaled ('See [tests.md](tests.md)... [mocking.md](mocking.md)'), but the bundle contains no such files — no references/, scripts/, or assets/ directories exist — so both links are dangling and navigation breaks. Scored against the actual (empty) bundle structure, this is 'structure present but references not resolvable', between the broken/minimal-structure anchors and the good-structure anchor 4.

3 / 5

Total

15

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, well-structured description with an explicit and multi-trigger 'Use when...' clause in third person. Its main weakness is that the 'what' is limited to the domain name 'Test-driven development' without naming the skill's concrete capabilities, and a few natural synonyms (TDD, unit tests) are absent.

Suggestions

Extend the 'what' with 2-3 concrete capabilities, e.g., 'Test-driven development: agree test seams, write failing tests first, apply minimal implementations, and avoid tautological or implementation-coupled tests.'

Add common synonyms users say, such as 'TDD', 'write the tests first', or 'unit tests', to the trigger clause.

DimensionReasoningScore

Specificity

The description names the domain ('Test-driven development') and conveys the concrete activity ('build features or fix bugs test-first'), but lists no specific capabilities of the skill itself (e.g., test design guidance, seam selection, anti-patterns), so coverage is not comprehensive. It sits between anchor 2 (domain only, generic actions) and anchor 4 (several specific actions) — it is more than a bare domain label but lacks the skill-side action list of anchor 4.

3 / 5

Completeness

Both parts exist: 'what' ('Test-driven development') and an explicit 'when' clause with multiple concrete triggers. It falls short of anchor 5 only because the 'what' is a bare domain name without stating what the skill actually provides (test-quality rules, seam agreement, anti-patterns), leaving the capability side thinner than the trigger side.

4 / 5

Trigger Term Quality

Natural trigger phrases are present: 'build features', 'fix bugs test-first', 'red-green-refactor', 'integration tests' — terms a user would plausibly say. A few common variations are missing (e.g., 'TDD', 'write tests first', 'unit tests'), matching the 'good keyword coverage; a few natural terms missing' anchor rather than the comprehensive synonym/extension coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

TDD is a clear niche with distinct triggers ('test-first', 'red-green-refactor'), so it is mostly distinguishable from general coding or code-review skills. Minor overlap risk remains: 'wants integration tests' alone could pull this skill in when the user just wants tests written, not a TDD loop — matching 'mostly distinct; minor overlap risk' rather than the minimal-conflict anchor 5.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing

Warning

Total

15

/

16

Passed

Repository
MODSetter/SurfSense
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.