CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-tdd

Build a behavior change with observed red, minimal green, and measured test consolidation

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-tdd/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, highly disciplined TDD instruction set with an explicit hard gate, sequenced steps, feedback loops for both failing tests and stalled strategies, and unusually concrete validation criteria (mutant killing, measured medians, 20%/100ms thresholds). Its only notable weakness is that the referenced plugin files are not part of this skill's bundle, leaving those navigation pointers unresolvable locally.

DimensionReasoningScore

Conciseness

The ~55-line body is dense and assumes competence throughout — no explanation of what TDD is, no padding, every rule stated once. Phrases like 'If a test passes before implementation, it is not red evidence' carry maximal instruction per token, matching the 'lean and efficient; every token earns its place' anchor.

5 / 5

Actionability

Guidance is concrete for an instruction-only skill: a six-step numbered rule, an explicit ledger schema ('old_test', 'behavior', 'replacement', 'mutant', 'red_observed', 'baseline_ms', 'candidate_ms', 'reason'), and measurable thresholds ('five isolated runs', '20 percent and 100 ms'). Minor gaps remain — no example ledger entry or sample commands for the measurement step — keeping it below fully copy-paste-ready.

4 / 5

Workflow Clarity

The red/green/refactor sequence is explicitly numbered with validation at every step ('Run it and record the expected failure'), error-recovery loops ('repair the test until it fails on the missing behavior', strategy rotation after two failed attempts), and a completion checklist (red/green evidence, affected-suite results, consolidation ledger). The destructive test-removal workflow is well guarded by the mutant-killing requirement, so the batch/destructive cap does not apply.

5 / 5

Progressive Disclosure

Sections are well organized for a skill under 50 lines, and external references ('skills/blocks/engineering-method-selection.md', 'skills/blocks/codex-host-adapter.md', 'THIRD_PARTY_NOTICES.md') are clearly signaled with stated purpose and one level deep. However, no bundle files exist alongside SKILL.md — the referenced paths live outside this skill's directory, a minor navigation gap that keeps it just below a 5.

4 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and third-person with a clear articulation of what the skill does, but it omits any 'when to use' trigger guidance and lacks the natural trigger terms ('TDD', 'test-driven development', 'red-green-refactor') a user would actually say. It is functional but below the standard of the good examples, which pair concrete actions with explicit use-when clauses.

Suggestions

Add an explicit trigger clause, e.g. 'Use when asked to implement a behavior change test-first, practice TDD, or when the user mentions red-green-refactor or writing tests before implementation.'

Include the natural synonyms users say — 'test-driven development', 'TDD', 'red-green-refactor', 'write tests first' — alongside the current red/green terminology.

State the when-condition context for the consolidation half of the skill (e.g. 'Use when consolidating or deduplicating a test suite') so both capabilities have trigger coverage.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'observed red', 'minimal green', 'measured test consolidation' — naming the TDD mechanics precisely rather than vaguely. It stops short of comprehensive coverage (no explicit refactor/testing-suite mention), matching the 'several specific actions; minor gaps' anchor.

4 / 5

Completeness

The 'what' is clear (build a behavior change through red/green with test consolidation), but there is no 'Use when...' clause or equivalent trigger guidance, capping completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Relevant keywords like 'behavior change', 'test', 'red', and 'green' are present, but the natural phrases users actually say — 'test-driven development', 'TDD', 'red-green-refactor', 'write tests first' — are missing. This matches 'some relevant keywords but missing common variations or synonyms'.

3 / 5

Distinctiveness Conflict Risk

The red/green/consolidation framing carves out a distinct TDD niche unlikely to trigger general testing or refactoring skills. Minor overlap risk remains with plain test-writing skills, matching 'mostly distinct; minor overlap risk'.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.