CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-development

Drives development with tests using the red-green-refactor loop. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.

60

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/test-driven-development/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body excels at workflow clarity — the RED/GREEN/REFACTOR and Prove-It loops are explicitly sequenced with validation checkpoints, feedback loops, and a completion checklist — and its guidance is largely concrete with real code examples. Its two weaknesses are noticeable verbosity from re-teaching standard testing doctrine Claude already knows, and a monolithic structure whose two external references are vaguely signaled and not backed by any bundle files.

Suggestions

Cut or drastically compress sections restating knowledge Claude already has (Arrange-Act-Assert, one-assertion-per-concept, good/bad test naming, mocks-vs-fakes, the Common Rationalizations table) to a few lines of policy each.

Move the Browser Testing with DevTools section into a reference file (e.g., references/browser-testing.md) and link to it, keeping only the security-boundary warning inline.

Fix the dangling references: link browser-testing-with-devtools and testing-patterns.md with clear, valid paths inside the skill's own references/ directory so navigation is unambiguous.

DimensionReasoningScore

Conciseness

The ~400-line body re-teaches testing knowledge Claude already has at length: the Arrange-Act-Assert pattern, one-assertion-per-concept, descriptive test naming with good/bad examples, the mocks-vs-stubs-vs-fakes preference order, the test pyramid, a "Common Rationalizations" motivational table, and the "Beyonce Rule" quip — several padded sections of standard doctrine, matching "noticeably verbose; several unnecessary explanations or padded sections". It is more purposeful than the heavily padded level 1 (the Discover-the-Stack-First and Prove-It sections add genuine workflow instruction Claude might not follow unprompted), so it does not fall to 1; but the standard-testing-tutorial material keeps it clearly below the "minor instances" level 4.

2 / 5

Actionability

Guidance is mostly concrete and executable: runnable TypeScript test examples for each cycle step ("expect(task.status).toBe('pending')"), a worked bug-reproduction example, a concrete stack-discovery checklist ("package.json, pom.xml/build.gradle, pyproject.toml...") with named wrapper commands ("./gradlew", "make test"), a test-size decision guide, and an end-of-work verification checklist. Minor gaps keep it from 5: the code examples are illustrative fragments (taskService and db are never defined), and no commands are copy-paste ready since the skill deliberately defers to the repo's own tooling.

4 / 5

Workflow Clarity

The core loop is an explicit, clearly sequenced workflow with validation checkpoints at every step — RED requires the test to FAIL ("A test that passes immediately proves nothing"), GREEN requires it to pass, REFACTOR requires "Run tests after every refactor step to confirm nothing broke" — and the Prove-It pattern adds a full feedback loop ending in "Run full test suite (no regressions)". A closing checklist ("Every new behavior has a corresponding test", "Bug fixes include a reproduction test that failed before the fix") matches the top anchor: clear sequence, explicit validation, feedback loops, and checklists.

5 / 5

Progressive Disclosure

Section structure is good, but the skill is monolithic: no references/ bundle exists, yet the body inlines ~40 lines of browser/DevTools testing that clearly belong in a separate file, and it points to two external resources — the bare name "browser-testing-with-devtools" and a path outside the skill directory ("../../references/testing-patterns.md") — that are neither clearly signaled nor present as bundle files. This fits "some structure but could be better organized; references present but not clearly signaled; content that should be separate is inline"; it is well above the unstructured level 2 thanks to clean headers and a genuinely useful overview role.

3 / 5

Total

14

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly and explicitly states both what the skill does and when to use it, with natural, varied trigger phrases. Its weaknesses are mild: the capability statement is a single action rather than a list of concrete capabilities, and the deliberately broad triggers ("any logic", "any bug") create overlap risk with general coding and debugging skills.

Suggestions

Enumerate one or two more concrete capabilities alongside the loop (e.g., "reproduce bugs with failing tests before fixing them") to lift specificity beyond a single named method.

Narrow or sharpen the trigger clauses (e.g., "when writing new logic or receiving a bug report") so the description does not compete with every general coding request.

DimensionReasoningScore

Specificity

"Drives development with tests using the red-green-refactor loop" names the domain and exactly one concrete method (the red-green-refactor loop), matching the anchor for naming a domain with 1-2 concrete actions but not comprehensive coverage. It does not list several distinct capabilities (as in "Extract text and tables, fill forms, merge documents"), so score 4 is not warranted; it is far more concrete than the generic "Processes PDF files" level.

3 / 5

Completeness

It explicitly answers both questions: the "what" ("Drives development with tests using the red-green-refactor loop") and an explicit, multi-clause "when" ("Use when implementing any logic, fixing any bug, or changing any behavior... when a bug report arrives, or when you're about to modify existing functionality") with concrete trigger phrases. This matches the top anchor exactly; level 4 would require the "when" to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Natural trigger phrases are present and varied — "implementing any logic", "fixing any bug", "when a bug report arrives", "modify existing functionality", "prove that code works" — which users would plausibly say. A few common variations are missing (e.g., "write tests", "add tests", "unit tests", "test coverage", "refactor"), keeping it just below the comprehensive-synonym level 5.

4 / 5

Distinctiveness Conflict Risk

The methodology itself (red-green-refactor, test-first) is a distinct niche, but the trigger scope — "implementing any logic", "fixing any bug", "changing any behavior" — is extremely broad and would plausibly fire for nearly any coding task, overlapping with general coding, debugging, and verification skills. It is more specific than the "very broad, high overlap" level 2 because the TDD framing is unmistakable, but the breadth of triggers prevents the minor-overlap level 4.

3 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
addyosmani/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.