CtrlK
BlogDocsLog inGet started
Tessl Logo

lock-tests

Lock the full test inventory before any implementation code is written. Reads spec+plan+AC, writes ALL failing tests in a batch, emits a Test Inventory doc, and gates with user approval.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced workflow with explicit validation checkpoints and strong feedback loops around the risky batch and git operations. Its main weakness is token efficiency — the same wrapper-delegation fact is restated several times and some paragraphs are run-on prose that could be tightened without losing clarity.

Suggestions

State the devflow-wrapper delegation fact ('wrapper delegates to upstream executing-plans and forces the post-implementation handoff to /devflow:finish-feature') once, in the attribution block or Phase 2, and reference it tersely elsewhere instead of repeating the full explanation three to four times.

Break the Phase 2 run-on paragraph into short steps or a small list (commit artefacts → gate via AskUserQuestion → spawn session with embedded handoff context) so the handoff mechanics scan as quickly as the other phases.

Move the upstream TDD constraint table and rules carried from haletothewood into a brief pointer, keeping only the mock-boundaries and UI-query rules that this phase actually enforces.

DimensionReasoningScore

Conciseness

Mostly efficient operational content, but the devflow-wrapper-delegates-to-upstream explanation ('the devflow wrapper that delegates to upstream `superpowers:executing-plans` … forces the post-implementation handoff to `/devflow:finish-feature`') is repeated nearly verbatim in the attribution block, Phase 0 step 6, Phase 2, and Phase 3, and the Phase 2 paragraph is a single ~120-word run-on sentence that could be tightened.

3 / 5

Actionability

Fully executable throughout: copy-paste-ready bash for worktree recovery (including the awk fallback), exact AskUserQuestion questions and options, a complete Test Inventory markdown template, and framework-specific test-run commands covering jest/vitest/pytest/rspec.

5 / 5

Workflow Clarity

Clear phase sequence with explicit validation at every risky step: 'Each test MUST fail with… NOT acceptable: syntax errors, fixture-loading errors' with a fix-before-proceeding loop, worktree verification that exits 1 on branch mismatch, a dirty-tree prompt, and the mandatory Phase 1.8 approval gate — a full feedback loop around the batch test-write.

5 / 5

Progressive Disclosure

Single-file skill with no bundle files, well-organized phase headers and no nested references; the inline scripts and templates are operationally needed. Minor gaps: the attribution block and the repeated Phase 2/3 handoff prose could be trimmed or split out.

4 / 5

Total

17

/

20

Passed

Description

80%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, concrete, third-person description that clearly states what the skill does and when it applies. It would benefit from explicit trigger phrasing and common synonyms like 'TDD' or 'test-first' to strengthen discovery and distinction from general test-writing skills.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Lock the full test inventory', 'Reads spec+plan+AC', 'writes ALL failing tests in a batch', 'emits a Test Inventory doc', 'gates with user approval' — comprehensively covering the skill's behavior in third-person voice.

5 / 5

Completeness

The 'what' is explicit and detailed, and 'before any implementation code is written' answers 'when', but as pipeline timing rather than an explicit 'Use when…' trigger phrase — it could be more explicit about the invoking context.

4 / 5

Trigger Term Quality

Good natural keywords ('test inventory', 'failing tests', 'spec', 'plan', 'user approval') but misses common variations a user would say, such as 'TDD', 'test-first', 'red/green', or 'write tests'.

4 / 5

Distinctiveness Conflict Risk

Clear niche (pre-implementation test-locking gate) with distinct triggers, though it could overlap with generic TDD or test-writing skill requests.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
AndreJorgeLopes/devflow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.