CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-development

TDD: enforce RED-GREEN-REFACTOR, tests before code.

57

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/software-development/test-driven-development/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a highly actionable, well-sequenced TDD workflow with executable commands, worked examples, mandatory verification checkpoints, and recovery guidance. Its main weakness is redundancy — the rationalization-rebuttal content appears in three overlapping sections that could be consolidated or moved to a reference file.

DimensionReasoningScore

Conciseness

Sentence-level style is lean and imperative with no tutorial padding about concepts Claude already knows, but the same anti-rationalization content is covered three times ("Why Order Matters", the "Common Rationalizations" table, and "Red Flags"), and "Final Rule" restates "The Iron Law". This matches 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than the minor-trimming of score 4 or the padded verbosity of score 2.

3 / 5

Actionability

The body gives copy-paste-ready material throughout: exact pytest commands with test-path selectors, complete good/bad test examples, a filled-in delegate_task template, and concrete problem/solution tables. This matches the anchor for fully executable, specific examples covering the common cases; it is well above the 'concrete code with minor gaps' level of score 4.

5 / 5

Workflow Clarity

The RED-GREEN-REFACTOR cycle is sequenced with mandatory verification steps at each phase ("Verify RED — Watch It Fail", "Verify GREEN — Watch It Pass" plus a full-suite regression run), explicit error-recovery feedback loops ("Test fails? Fix the code, not the test", "If tests fail during refactor: Undo immediately"), and a completion checklist. This matches the top anchor: clear sequence, explicit validation, feedback loops, and a checklist.

5 / 5

Progressive Disclosure

The single SKILL.md has no bundle files, and the content is organized under clear, well-ordered sections with an overview up front, so structure is good. However, at roughly 350 lines, the rebuttal/reference material (rationalizations, red flags) is inlined where it could be split into a reference file, which keeps it at 'good structure with minor organization gaps' rather than the exemplary split of score 5.

4 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names its niche, but it omits any 'when to use' trigger guidance and relies on the abbreviation 'TDD' without spelling out 'test-driven development' or common user phrasings. It reads more as a rule statement than a capability-plus-trigger description.

Suggestions

Add an explicit trigger clause, e.g. "Use when writing new features, fixing bugs, or refactoring, or when the user asks for test-driven development, unit tests, or tests-first workflow."

Spell out "test-driven development" alongside "TDD" and include natural user phrasings like "write tests first" to improve trigger-term coverage.

Name one or two more concrete capabilities (e.g. "write failing tests, verify minimal implementation, refactor with the suite green") to raise specificity beyond the single core action.

DimensionReasoningScore

Specificity

The description names its domain ("TDD") and one core concrete action ("enforce RED-GREEN-REFACTOR, tests before code"), matching the anchor for 1-2 concrete actions without comprehensive coverage. It does not list several specific actions like the score-4 anchor, and it is more than the minimal/generic naming of the score-2 anchor.

3 / 5

Completeness

The 'what' is stated clearly (enforce the red-green-refactor cycle, tests before code), but there is no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Score 4 requires both 'what' and an explicit 'when', which is absent here.

3 / 5

Trigger Term Quality

It includes relevant keywords ("TDD", "tests", "RED-GREEN-REFACTOR") but misses common variations users would naturally say, such as "test-driven development" spelled out, "write tests first", or "unit tests". This matches 'some relevant keywords but missing common variations'; it lacks the fuller synonym coverage of score 4.

3 / 5

Distinctiveness Conflict Risk

"TDD" and "RED-GREEN-REFACTOR" mark a clear niche that is unlikely to trigger for unrelated skills, but the generic term "tests" carries minor overlap risk with closely related general-testing or debugging skills. It is mostly distinct with minor overlap (score 4) rather than the minimal-conflict clarity of score 5.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.