CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-development

Implements a feature or bug fix test-first - write one failing test, watch it fail, write the minimal code to pass, then refactor with tests green. This is the implementation stage inside the delivery-flow workflow, entered by requests like "implement this feature", "fix this bug", "write the code for this", or "add the tests and implement it" once delivery-flow is driving the change. Do not activate directly for a standalone request that has not gone through delivery-flow first.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured skill body with complete executable examples, explicit verification checkpoints, feedback loops, and a properly signaled one-level-deep reference. Its main weakness is redundancy — the Red Flags section, Iron Law/Final Rule, and parts of the Good Tests table duplicate other sections or the reference file, inflating token cost without adding guidance.

Suggestions

Merge the 'Red Flags - STOP and Start Over' list into the 'Common Rationalizations' table (or reduce it to a pointer), since nearly every entry duplicates a table row.

Cut 'The Iron Law' or 'Final Rule' section — they state the same rule twice; keep one authoritative statement.

Trim the multi-sentence 'Reality' column entries to one line each, or move the extended rebuttals into references/writing-good-tests.md, keeping the main file lean.

DimensionReasoningScore

Conciseness

The body is punchy and imperative, but contains notable redundancy: the 'Red Flags' list re-indexes entries already in the 'Common Rationalizations' table ('Keep as reference', 'Already spent X hours', 'manually tested', 'spirit not ritual' appear in both), 'Final Rule' restates 'The Iron Law', and the rationalization table entries are wordy. This fits the 3 anchor ('could be tightened'); it is above 2 because there is no padded explanation of concepts Claude already knows.

3 / 5

Actionability

Fully executable guidance throughout: complete TypeScript test and implementation examples with Good/Bad contrasts, concrete commands ('npm test path/to/test.test.ts'), and a worked bug-fix example showing actual FAIL/PASS output. Copy-paste ready examples covering the common cases match the 5 anchor.

5 / 5

Workflow Clarity

Red-Green-Refactor is clearly sequenced with mandatory verification checkpoints ('Confirm: Test fails (not errors)', 'Failure message is expected'), explicit feedback loops for error recovery ('Test errors? Fix error, re-run until it fails correctly', 'Test fails? Fix code, not test'), and a completion checklist. This matches the 5 anchor including validation steps and error-recovery loops.

5 / 5

Progressive Disclosure

One well-signaled, one-level-deep reference (references/writing-good-tests.md, verified to exist, with its own 'Load this reference when' header) plus a bullet preview of its contents and otherwise well-organized sections. Minor gaps keep it at 4: the 'Good Tests' table overlaps content in the reference file, and the bulky multi-sentence rationalization entries sit inline where they could be referenced out.

4 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly states concrete capabilities, gives natural trigger phrases for when to activate, and includes an explicit negative boundary to avoid mis-activation. Its only weaknesses are missing common synonyms (TDD, test-driven, write tests first) and trigger phrases that are individually generic, relying on workflow gating rather than distinctive keywords for separation.

Suggestions

Add common synonyms such as 'TDD', 'test-driven development', or 'write the tests first' to the trigger phrase list so the description matches how users naturally name the practice.

Make trigger terms more distinctive of the test-first approach (e.g. 'add the tests and implement it') rather than generic implementation requests, reducing reliance on the delivery-flow gate alone.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions covering the full cycle — 'write one failing test, watch it fail, write the minimal code to pass, then refactor with tests green' — with comprehensive coverage of the test-first workflow. It matches the 5 anchor rather than 4 because no meaningful action in the cycle is omitted.

5 / 5

Completeness

Explicitly answers both what ('Implements a feature or bug fix test-first - write one failing test, watch it fail...') and when ('entered by requests like "implement this feature", "fix this bug"... once delivery-flow is driving the change') with concrete trigger phrases, plus a negative activation boundary. This matches the 5 anchor; a 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Includes natural request phrases users would actually say ('implement this feature', 'fix this bug', 'write the code for this', 'add the tests and implement it'), but misses common synonyms such as 'TDD', 'test-driven', or 'write the tests first'. Good coverage with a few natural terms missing fits the 4 anchor, below 5 which requires synonym-level comprehensiveness.

4 / 5

Distinctiveness Conflict Risk

The delivery-flow gating ('Do not activate directly for a standalone request that has not gone through delivery-flow first') carves a clear niche, but the trigger phrases themselves ('fix this bug', 'implement this feature') are generic and overlap with any general implementation skill. Mostly distinct with minor overlap risk fits the 4 anchor; a 5 would require the triggers themselves to be distinctive.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tesslio/tessl-eval-demo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.