CtrlK
BlogDocsLog inGet started
Tessl Logo

verification-before-completion

Internal verification gate for delivery-flow. Use before delivery-flow claims work is complete, fixed, or passing; do not activate directly for a standalone user request.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/sdlc-assurance/skills/verification-before-completion/SKILL.md

The canonical home for this skill is verification-before-completion in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a focused, well-structured behavioral skill with a genuinely actionable gate procedure and strong validation thinking. Its weaknesses are redundancy across the red-flags, rationalization, and when-to-apply sections and placeholder-level rather than fully concrete command examples.

Suggestions

Merge "Red Flags - STOP", "Rationalization Prevention", and "When To Apply" into a single section — "just this once", "should", and trusting agent reports each currently appear two to three times.

Add one fully concrete worked example (e.g. "npm test → 34/34 pass → 'All tests pass'") to replace the generic "[Run test command]" placeholders.

Add an explicit fix-and-re-verify loop to the Gate Function (when verification fails: fix, re-run, re-read) to close the feedback-loop gap.

DimensionReasoningScore

Conciseness

The body is mostly lean ("NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE", terse tables) and assumes Claude's competence, but "Red Flags - STOP", "Rationalization Prevention", and "When To Apply" repeat the same items ("just this once", "should", trusting agent reports each appear multiple times), and the line "Violating the letter of this rule is violating the spirit of this rule" is cryptic filler. More than minor trimming is needed, so it sits below anchor 4.

3 / 5

Actionability

The Gate Function ("IDENTIFY: What command proves this claim? / RUN ... / READ: Full output, check exit code, count failures / VERIFY") and the ✅/❌ pattern pairs give mostly executable behavioral guidance, appropriate for an instruction-only skill. Not a 5 because commands remain generic placeholders ("[Run test command]") with no worked example naming a real command.

4 / 5

Workflow Clarity

The Gate Function is a clearly sequenced 5-step procedure with an explicit validation branch ("If NO: State actual status with evidence") and the regression-test pattern includes a red-green cycle ("Revert fix → Run (MUST FAIL) → Restore → Run (pass)"). Not a 5 because some checkpoints are implicit — there is no fix-and-re-verify loop after a failed verification step.

4 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent) and none are needed; the body uses clear section headers and tables for navigation. It misses anchor 5 because at ~115 lines it exceeds the simple-skill threshold and contains duplicated inline material that could be consolidated.

4 / 5

Total

15

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is short, third-person, and honest about its internal scope, with explicit when-guidance and strong distinctiveness. Its main weaknesses are the abstract "verification gate" framing with no concrete actions and thin trigger-term coverage.

Suggestions

Replace the abstract "verification gate" framing with 1-2 concrete actions, e.g. "Runs verification commands and reads their output before allowing delivery-flow to claim work is complete, fixed, or passing."

Broaden trigger coverage with natural synonyms such as "done", "works", "all tests green", or "ready to ship" so the skill activates across real variations of completion claims.

State what the skill actually does on activation (e.g. apply the identify-run-read-verify gate) so the 'what' half is as explicit as the 'when' half.

DimensionReasoningScore

Specificity

"Internal verification gate for delivery-flow" names the domain but lists no concrete actions — there is no equivalent of "runs tests / checks exit codes", only the abstract noun "verification gate". It does not reach anchor 3 because no 1-2 concrete actions are stated.

2 / 5

Completeness

Both parts are present: what ("Internal verification gate for delivery-flow") and when ("Use before delivery-flow claims work is complete, fixed, or passing"), plus an explicit exclusion ("do not activate directly"). Not a 5 because the trigger guidance covers a narrow set of phrases and could be more explicit about the full range of completion claims.

4 / 5

Trigger Term Quality

"work is complete, fixed, or passing" provides some natural, relevant trigger terms, but common variations and synonyms (e.g. "done", "works", "all green", "ship") are missing. It is above anchor 2 because the terms present are domain-appropriate rather than generic.

3 / 5

Distinctiveness Conflict Risk

The description is explicitly scoped as "Internal ... for delivery-flow" with "do not activate directly for a standalone user request", giving it a clear niche with minimal conflict risk against user-facing skills. Nothing overlaps with generic skill triggers.

5 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tesslio/tessl-eval-demo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.