CtrlK
BlogDocsLog inGet started
Tessl Logo

verification-before-completion

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always

80

1.22x
Quality

72%

Does it follow best practices?

Impact

92%

1.22x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./eval/local/skills/benchmarks/dependency/superpowers/verification-before-completion/SKILL.md

The canonical home for this skill is verification-before-completion in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

66%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill delivers a crisp, validated workflow — the Gate Function with its explicit verify branch and the red-green regression pattern are exemplary — supported by concrete claim-to-evidence mappings. Its main weakness is redundancy: the same injunction is repeated across four sections and two overlapping tables, which wastes tokens without adding guidance.

Suggestions

Collapse the redundant restatements: merge 'The Iron Law' and 'The Bottom Line' into the Overview's core principle, keeping the Gate Function as the single detailed statement of the rule.

Merge the 'Red Flags - STOP' and 'Rationalization Prevention' tables into one 'Excuse → Required Action' table, since entries like 'Trusting agent success reports' / 'Agent said success' and 'Just this once' appear in both.

Trim 'Why This Matters' to the one-line trust consequence and drop the enumerated failure-memory list, which restates cases already covered by the Common Failures table.

DimensionReasoningScore

Conciseness

The core rule is restated at least four times ("Evidence before claims, always", "NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE", the Gate Function's "Skip any step = lying", and "Run the command. Read the output. THEN claim the result"), and the "Red Flags" and "Rationalization Prevention" tables substantially duplicate each other. This matches anchor 2's 'several unnecessary explanations or padded sections'; it is not 1 because no space is spent explaining concepts Claude already knows.

2 / 5

Actionability

The Gate Function gives a concrete executable procedure ("IDENTIFY: What command proves this claim? ... RUN: Execute the FULL command ... READ: Full output, check exit code, count failures") and the Common Failures table maps specific claims to required evidence ("Tests pass | Test command output: 0 failures"). It is not 5 because guidance uses placeholders like "[Run test command]" rather than copy-paste-ready commands, though for a domain-agnostic behavioral skill those placeholders are defensible; it clearly exceeds anchor 3's pseudocode-level guidance.

4 / 5

Workflow Clarity

The Gate Function is a clearly sequenced 5-step workflow with an explicit validation checkpoint and branching ("VERIFY: Does output confirm the claim? - If NO: State actual status with evidence - If YES: State claim WITH evidence"), and the regression-test pattern encodes a full red-green feedback loop ("Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)"). This matches anchor 5's explicit validation steps with error-recovery feedback loops.

5 / 5

Progressive Disclosure

The body is well organized into clearly labeled, scannable sections (Overview, Iron Law, Gate Function, Common Failures, Key Patterns, When To Apply) and needs no external references — no references/, scripts/, or assets/ exist. It does not reach 5 because at ~135 lines with heavily overlapping sections, tighter structure or splitting the rationale/examples would improve navigation; it is better than anchor 3 since nothing is buried or misfiled.

4 / 5

Total

15

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong behavioral-skill description: it explicitly couples a clear 'what' (run verification and confirm output before success claims) with a concrete 'Use when' trigger clause covering completion claims, commits, and PRs. Trigger terms are natural and the niche is distinct, though specificity of listed actions and synonym coverage leave small room for improvement.

DimensionReasoningScore

Specificity

The description names its domain and a couple of concrete actions ("requires running verification commands and confirming output before making any success claims", "before committing or creating PRs"), but does not enumerate several specific verification behaviors. It is not 4 because the action list is minimal rather than 'several specific actions with minor gaps', and not 2 because it goes beyond naming the domain with generic language.

3 / 5

Completeness

Both halves are explicit: 'when' via "Use when about to claim work is complete, fixed, or passing, before committing or creating PRs" and 'what' via "requires running verification commands and confirming output before making any success claims". This matches anchor 5 (clear, explicit what AND when with concrete trigger phrases); score 4 would require the 'when' to be less explicit than it is.

5 / 5

Trigger Term Quality

Natural phrases users/sessions would produce are present: "claim work is complete, fixed, or passing", "committing", "creating PRs". It falls short of anchor 5 because common synonyms and variations like "done", "tests pass", "it works", or "verify before claiming" are missing, but it clearly exceeds anchor 3's 'some relevant keywords with missing variations'.

4 / 5

Distinctiveness Conflict Risk

It occupies a clear niche (verification discipline before completion claims) with triggers unlikely to fire for unrelated skills, but phrases like "fixed" or "passing" have minor overlap risk with testing/review-type skills. It does not reach anchor 5's 'minimal conflict risk' but is more distinct than anchor 3's 'could still overlap with similar skills'.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
rpamis/comet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.