CtrlK
BlogDocsLog inGet started
Tessl Logo

behavioural-tdd

Execute a strict Red-Green-Refactor TDD cycle for one requirement at a time while applying KISS and DRY. Use when the user provides a business rule, acceptance criterion, bug fix, or feature requirement and wants test-first development, behavioral tests, failing tests first, TDD, red-green-refactor, or high-quality code. Works for unit, integration, UI component, API, and CLI tests across stacks.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, high-quality process skill: lean prescriptive rules with no filler, concrete per-phase steps and required outputs, and appropriately split detail across two existing one-level-deep reference files. The single gap is the absence of an explicit error-recovery loop (what to do when GREEN or REFACTOR validation fails) that would complete the workflow's feedback cycle.

DimensionReasoningScore

Conciseness

The body is lean and purely prescriptive — "Execute one complete TDD cycle per requirement. Never test private internals." — with zero space spent explaining what TDD, mocking, or KISS/DRY are. Every section adds rules Claude does not already know (e.g., "Shameless Green is acceptable for a single new test"), matching the 5 anchor; the only near-redundancy (two overlapping UI-query bullets) is too minor to drop it to the 4 anchor.

5 / 5

Actionability

For an instruction-only skill the guidance is fully executable: each phase has an Input, numbered steps, and explicit present-items, including a copy-paste-ready literal ("The behavioral test `<test name>` still passes.") and concrete selector commands ("Prefer `getByRole` with accessible name"), with examples delegated to references/tests.md. This matches the 5 anchor's specificity and coverage; the 4 anchor would require missing key execution details, which are absent.

5 / 5

Workflow Clarity

The RED→GREEN→REFACTOR sequence is explicit with inputs, outputs, and validation checkpoints ("Confirmation that the test now passes, or the exact blocker if it cannot be run", the stop-gate "Stop after RED and wait for confirmation", and a Constraints checklist). It falls short of the 5 anchor because there is no error-recovery loop — e.g., what to do when the refactor breaks the test — only blocker reporting, which fits the 4 anchor (clear sequence, most checkpoints, minor gaps).

4 / 5

Progressive Disclosure

Clear overview sections with two well-signaled, contextual, one-level-deep references — "Read `references/mocking.md` when dependency boundaries or mocks are involved" and "Read `references/tests.md` for examples and red flags" — both verified to exist as real, substantive files (65 and 58 lines), with no nesting. This matches the 5 anchor exactly; the 4 anchor would require buried references or inline content that belongs in a separate file, which is not the case.

5 / 5

Total

19

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, explicit what-and-when structure, and rich natural trigger vocabulary covering TDD synonyms and requirement types. The only weaknesses are a modestly broad pair of triggers ("bug fix", "high-quality code") that could over-fire, and a what-clause that describes one process rather than a set of discrete actions.

DimensionReasoningScore

Specificity

Names several concrete actions — "Execute a strict Red-Green-Refactor TDD cycle for one requirement at a time while applying KISS and DRY" plus "Works for unit, integration, UI component, API, and CLI tests across stacks" — but the core action is one process plus principles rather than the multiple distinct actions of the 5 anchor (e.g., extract, fill, merge, convert). Fits the 4 anchor (several specific actions, minor gaps) better than 3, since it goes well beyond naming the domain with 1-2 generic actions.

4 / 5

Completeness

Explicitly answers both: the what is "Execute a strict Red-Green-Refactor TDD cycle for one requirement at a time while applying KISS and DRY" and the when is "Use when the user provides a business rule, acceptance criterion, bug fix, or feature requirement and wants test-first development..." — concrete trigger phrases for both, matching the 5 anchor exactly. Clearly above the 4 anchor, where the when would be less specific.

5 / 5

Trigger Term Quality

Comprehensive natural terms with synonyms: "test-first development, behavioral tests, failing tests first, TDD, red-green-refactor" alongside "business rule, acceptance criterion, bug fix, or feature requirement" — all phrases a user would naturally say when needing this skill. Not the 4 anchor because no common variation of the trigger vocabulary is missing; file extensions are inapplicable to a process skill.

5 / 5

Distinctiveness Conflict Risk

TDD with strict red-green-refactor triggers is a clear niche with mostly distinct triggers, but "bug fix" and "high-quality code" are broad enough to fire on ordinary bug-fix or code-quality requests where the user did not ask for test-first development — minor overlap risk with closely related coding skills, matching the 4 anchor rather than the minimal-conflict 5 anchor.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
pstember/codex
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.