CtrlK
BlogDocsLog inGet started
Tessl Logo

test-driven-development

Follow TalkPipe TDD workflow when making code changes. Use when implementing features, fixing bugs, or refactoring. Enforces write-fail-fix-pass order and test placement conventions.

86

0.83x
Quality

89%

Does it follow best practices?

Impact

67%

0.83x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exemplary: a strictly ordered workflow with explicit verify-fail and verify-pass checkpoints, concrete pytest commands, environment conventions, an explicit fallback policy, and a closing checklist, all with zero padding. Nothing in the rubric's failure patterns is present.

DimensionReasoningScore

Conciseness

The body never explains what TDD is or pads with background; every section (workflow, conventions, commands, fallback policy, checklist) carries operational weight. Matches the anchor for lean, efficient content that assumes Claude's competence. Not 4 because there is no over-explanation to trim.

5 / 5

Actionability

Commands are copy-paste executable throughout: "pytest --cov=src", "pytest tests/test_foo.py::test_bar -v", "source .venv/bin/activate", covering all common run modes. Matches the anchor for fully executable guidance covering common cases. Not 4 because there are no gaps; the only placeholder, "pytest tests/... -v", is immediately backed by exact examples in the Commands section.

5 / 5

Workflow Clarity

The "Workflow (Strict Order)" section sequences write-fail-change-pass with explicit validation checkpoints ("Verify the test fails", "Verify the test passes"), an error-recovery path for when a test is not reasonable ("Do not skip the test silently; explain why and get user confirmation"), and a closing checklist. Matches the anchor for clear sequence, explicit validation, feedback loops, and checklists. Not 4 because no checkpoint is missing or implicit.

5 / 5

Progressive Disclosure

This is a single-purpose, well-sectioned skill with no external references and no bundle files, so well-organized sections fully satisfy the rubric's simple-skill exception. Matches the anchor for a clear, easy-to-navigate structure. Not 4 because nothing is misplaced and no content needs splitting into separate files.

5 / 5

Total

20

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description cleanly answers both what and when with explicit trigger phrases and a distinct workflow identity. Its main weakness is specificity: it summarizes the workflow in one compressed clause rather than naming the concrete actions it performs.

Suggestions

Replace the compressed "Enforces write-fail-fix-pass order and test placement conventions" with named actions, e.g. "Writes failing unit tests first, verifies they fail, makes the change, then verifies they pass under pytest with coverage."

Add natural test-related trigger terms such as "writing tests", "unit tests", or "test coverage" to improve keyword coverage for users who ask for tests directly rather than framing it as a code change.

Consider narrowing the when-clause (e.g. "when changing Python source in this repo") to reduce overlap with generic coding and testing skills.

DimensionReasoningScore

Specificity

Phrases like "Enforces write-fail-fix-pass order and test placement conventions" name the domain and 1-2 concrete behaviors, but the description does not list several specific actions (no mention of writing tests, verifying failure, or running coverage). Not 4 because action coverage is narrow; not 2 because it goes beyond bare domain-naming.

3 / 5

Completeness

Both what ("Follow TalkPipe TDD workflow... Enforces write-fail-fix-pass order and test placement conventions") and when ("Use when implementing features, fixing bugs, or refactoring") are explicit and concrete, matching the anchor for clear answers to both questions. Not 4 because the when-clause is fully explicit rather than merely adequate.

5 / 5

Trigger Term Quality

"Use when implementing features, fixing bugs, or refactoring" provides natural trigger phrases users would actually say. Not 5 because common test-related terms ("writing tests", "test coverage", "unit tests") and their variations are missing; not 3 because several natural terms are present, not just one or two generic ones.

4 / 5

Distinctiveness Conflict Risk

"TalkPipe TDD workflow" and "write-fail-fix-pass order" carve a distinct niche, but the when-clause spans nearly all code changes, creating minor overlap risk with generic coding or testing skills. Not 5 because of that breadth; not 3 because the workflow identity is specific enough to mostly disambiguate.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
sandialabs/talkpipe
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.