CtrlK
BlogDocsLog inGet started
Tessl Logo

common-tdd

Guides quality-first TDD for new behavior, bug fixes, and test changes. Selects the smallest test layer, proves a distinct regression risk, and runs bounded RED-GREEN-REFACTOR verification.

58

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.github/skills/common/common-tdd/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured, lean, and features an excellent bounded workflow with explicit validation checkpoints and feedback loops. Its main weakness is progressive disclosure: only one of eight bundle files is signaled from the body.

Suggestions

Link the remaining relevant bundle files from the body (e.g. best-practices.md, tdd_patterns.md, test_runners.md, anti-patterns.md) with one-line 'See X for Y' pointers so Claude can discover them.

Consolidate the seven unreferenced references or move their contents into quality-contract.md so there is a single clearly-signaled deep reference rather than orphaned files.

Tighten the 'Red flags and rationalizations' prose into a compact stop-on list to shave a few tokens.

DimensionReasoningScore

Conciseness

The body is lean and avoids explaining concepts Claude already knows, with only minor instances of prose that could be tightened further (e.g. the 'Red flags and rationalizations' lead-in).

4 / 5

Actionability

Provides concrete, specific guidance for an instruction-only skill — the Test Intent Record fields (contract, fault, layer, cases, command) and the numbered bounded loop — though it stops short of literal copy-paste commands, which it deliberately derives per project.

4 / 5

Workflow Clarity

The numbered 'Bounded loop' has an explicit sequence with validation checkpoints (RED classification into expected_red/invalid_red/unexpected_green/verification_infra_failed), feedback loops (unexpected_green -> inspect and remove a redundant case), and a checklist (intent record fields, stop-on red flags).

5 / 5

Progressive Disclosure

The body is well-sectioned and signals one real one-level-deep reference ('See references/quality-contract.md'), but seven other bundle files in references/ (best-practices, tdd_patterns, test_runners, anti-patterns, etc.) are never referenced, leaving a navigation gap.

3 / 5

Total

16

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and targets a distinct TDD niche, but it omits an explicit 'Use when...' trigger clause and is missing common user phrasings like 'unit test' and 'failing test'. Completeness is consequently capped at 3.

Suggestions

Add an explicit 'Use when...' clause naming concrete triggers, e.g. 'Use when writing tests, doing TDD, fixing bugs with a regression test, or when the user mentions unit tests, failing tests, or red-green-refactor.'

Include common natural synonyms in the description such as 'unit test', 'write test', 'failing test', and 'test coverage' so it matches phrases users actually say.

Add file-extension or framework-agnostic trigger hints (e.g. 'new behavior, bug fixes, and test changes') to broaden distinctiveness without losing the TDD niche.

DimensionReasoningScore

Specificity

Lists several concrete TDD-specific actions ('Selects the smallest test layer', 'proves a distinct regression risk', 'runs bounded RED-GREEN-REFACTOR verification'), though the actions are slightly more abstract than the score-5 anchor.

4 / 5

Completeness

Provides a clear 'what' (guides quality-first TDD, selects layer, proves regression risk, runs RED-GREEN-REFACTOR) but lacks any explicit 'Use when...' or 'when should Claude use it' clause, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Contains relevant natural terms ('TDD', 'test', 'bug fixes', 'RED-GREEN-REFACTOR', 'regression') but misses common variations a user would say such as 'unit test', 'write test', 'failing test', or 'test coverage'.

3 / 5

Distinctiveness Conflict Risk

Targets a clear niche (quality-first TDD with RED-GREEN-REFACTOR) with distinct triggers, with only minor overlap risk against a more general testing skill.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
HoangNguyen0403/agent-skills-standard
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.