CtrlK
BlogDocsLog inGet started
Tessl Logo

tdd

Enforces strict Test-Driven Development with RED-GREEN-REFACTOR cycles. Writes one failing test at a time, implements minimal code to pass, then refactors. Delegates to the `test-provenance-guard` skill during REFACTOR to detect tests-by-construction (static + mutation checks). Pairs with the `code-quality` skill: invokes `Skill('code-quality')` during the REFACTOR phase to apply the full code-quality rule set against the GREEN output, and cites refactor recipes (R1–R20) by ID when reporting changes. Triggers on: "tdd", "write tests", "test this", "add test coverage", "test driven", "red green refactor", "/tdd".

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-crafted workflow skill: tight imperative rules, executable specifics (globs, naming examples, mocking decision rules), and a fully sequenced RED-GREEN-REFACTOR procedure with validation and error-recovery checkpoints. The only deductions are repeated code-quality routing explanations and references to `rules/*.md` files that are not present in the bundle for verification.

Suggestions

Consolidate the three explanations of Skill('code-quality') routing (Step 2 REFACTOR, GREEN phase note, and Code Quality section) into one authoritative statement to cut redundancy.

Ship the referenced `rules/red.md`, `rules/green.md`, `rules/refactor.md`, and `rules/test-after.md` files in the bundle so the clearly signaled one-level references resolve.

The GREEN-phase readability primitives duplicate content implied by the code-quality skill; a one-line pointer would suffice and reduce token cost.

DimensionReasoningScore

Conciseness

The body is dense, imperative rule-setting with almost no tutorial padding ("Write exactly ONE failing test. Run it. Confirm it fails with the expected error. Do NOT write implementation code."). Not 5 because the `Skill('code-quality')` routing rationale is explained three times (Step 2 REFACTOR, GREEN phase, Code Quality section) and could be consolidated; not 3 because nothing explains concepts Claude doesn't need explained.

4 / 5

Actionability

Guidance is copy-paste executable: exact glob patterns ("**/*.test.*", "**/*.spec.*"), runner discovery sources (package.json, Makefile, pyproject.toml...), a concrete naming example ("describe('createOrder')" → "it('should reject order when inventory is zero')"), factory pattern (`buildUser(overrides?)`), and a 3-mock design rule. Per the rubric's instruction-skill note, absence of code isn't penalized when guidance is this actionable — and it covers the common cases.

5 / 5

Workflow Clarity

Steps 0–4 are clearly sequenced with explicit validation checkpoints at every phase: RED ("Confirm it fails with the expected error"), GREEN ("Run the full relevant test suite to check for regressions"), completion ("stop and fix it before proceeding. Never accumulate broken tests"), plus a "When Things Go Wrong" section that gives error-recovery feedback loops and a 2-failed-attempts restart rule. This matches the strongest anchor's validate→fix→retry pattern.

5 / 5

Progressive Disclosure

The overview is well structured and phase details are split into clearly signaled one-level references ("See `rules/red.md`", "See `rules/green.md`", "See `rules/refactor.md`", "See `rules/test-after.md`"). Not 5 because no bundle directories (references/, scripts/, assets/, or rules/) are present, so the referenced rule files are dangling and navigation cannot be confirmed against the actual bundle; not 3 because the references that exist in the text are clearly signaled and content placement is appropriate.

4 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person capability statements, explicit trigger phrases, and clear cross-skill delegation behavior. Trigger coverage and distinctiveness are very good but miss a few common synonyms (e.g., "unit tests") that slightly broaden overlap with generic testing skills.

Suggestions

Add natural synonyms such as "unit tests" / "write unit tests" to the trigger list to capture the most common phrasing users would say.

Consider tightening the code-quality pairing sentence; the recipe-ID citation detail is more at home in the skill body than in the description.

DimensionReasoningScore

Specificity

Quotes like "Writes one failing test at a time, implements minimal code to pass, then refactors", "Delegates to the `test-provenance-guard` skill during REFACTOR to detect tests-by-construction (static + mutation checks)", and "cites refactor recipes (R1–R20) by ID" list multiple specific concrete actions with comprehensive coverage of the TDD workflow. Not 4 because coverage is thorough — the delegation and pairing behaviors are stated as concrete actions, not left as minor gaps.

5 / 5

Completeness

Both questions are answered explicitly: what — "Enforces strict Test-Driven Development with RED-GREEN-REFACTOR cycles... implements minimal code to pass, then refactors"; when — "Triggers on: 'tdd', 'write tests', 'test this', 'add test coverage', 'test driven', 'red green refactor', '/tdd'". The explicit trigger list meets the strongest anchor; no cap applies since trigger guidance is present.

5 / 5

Trigger Term Quality

"tdd", "write tests", "test this", "add test coverage", "test driven", "red green refactor", "/tdd" give good natural-phrase coverage. Not 5 because common synonyms like "unit tests", "test suite", or "testing" are missing; not 3 because the listed phrases are exactly what users naturally say and include slash-command form.

4 / 5

Distinctiveness Conflict Risk

"tdd", "test driven", "red green refactor" carve a clear niche distinct from generic testing skills, but "write tests" and "test this" could plausibly trigger for a general test-writing or coverage skill — minor overlap risk. Not 5 because of that broad-phrase overlap; not 3 because the core triggers are unmistakably TDD-specific.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.