CtrlK
BlogDocsLog inGet started
Tessl Logo

verification-before-completion

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is verification-before-completion in obra/superpowers

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable discipline skill: the gate function, evidence-mapping tables, and ✅/❌ pattern templates give Claude an exact procedure with validation and feedback loops built in. The only notable weakness is redundancy between the Red Flags and Rationalization tables plus one garbled sentence in the Overview.

Suggestions

Rewrite or delete the unclear Overview line 'Violating the letter of this rule is violating the spirit of this rule.' — as written it is garbled and adds no actionable content.

Merge the 'Red Flags - STOP' list and the 'Rationalization Prevention' table into a single excuse-vs-required-evidence table; they largely restate each other and both exist to catch the same drift.

Trim 'When To Apply', which restates the description's trigger conditions nearly verbatim; one cross-reference to the trigger set would reclaim those tokens.

DimensionReasoningScore

Conciseness

The body is lean for a discipline skill: tables, a code-block gate function, and terse ✅/❌ pattern templates with no padding or explanations of concepts Claude already knows. It stops short of anchor 5 because of real redundancy — the 'Red Flags' list and 'Rationalization Prevention' table cover overlapping ground, 'When To Apply' restates the description's triggers, and the garbled line 'Violating the letter of this rule is violating the spirit of this rule' earns no token.

4 / 5

Actionability

The Gate Function is a fully executable five-step procedure (IDENTIFY the command, RUN it fresh, READ output/exit code/failure count, VERIFY, only then claim), the claims table maps each assertion to its required evidence ('Tests pass' → 'Test command output: 0 failures'), and the Key Patterns are copy-paste-ready claim templates ('[Run test command] [See: 34/34 pass] "All tests pass"'). Per the rubric's instruction-skill note, the absence of literal code is not penalized when the guidance is this concrete.

5 / 5

Workflow Clarity

The gate is a clearly sequenced workflow with validation as its core: step 4 branches on evidence ('If NO: State actual status with evidence; If YES: State claim WITH evidence'), and the regression-test pattern embeds an explicit red-green feedback loop ('Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)'). This matches the anchor-5 example's structure of sequence, validation, and error-recovery loop.

5 / 5

Progressive Disclosure

This is a single-purpose, self-contained skill with no bundle files (no references/, scripts/, or assets/ exist), and everything present belongs inline — a gate procedure, two lookup tables, and pattern templates. Per the rubric's simple-skill guidance, well-organized sections with no need for external references score 5; sections are clearly headed and navigable with no monolithic wall or buried references.

5 / 5

Total

19

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' trigger clause, a concrete statement of required behavior, and natural trigger vocabulary. Minor room to improve: enumerate the kinds of verification in scope and add a few more colloquial completion phrases as triggers.

DimensionReasoningScore

Specificity

The description lists several concrete behaviors — 'requires running verification commands and confirming output before making any success claims' and 'evidence before assertions' — tied to specific contexts (complete, fixed, passing, committing, PRs). It falls below anchor 5 only because it never specifies which kinds of verification (tests, builds, linters) are in scope, leaving a minor coverage gap.

4 / 5

Completeness

Both halves are explicit and concrete: the 'what' ('requires running verification commands and confirming output before making any success claims') and the 'when' ('Use when about to claim work is complete, fixed, or passing, before committing or creating PRs'). This matches the anchor-5 example's structure of concrete capability statement plus explicit trigger phrases; there is no missing half to push it down.

5 / 5

Trigger Term Quality

Natural phrases a user or Claude would actually say are well covered: 'complete', 'fixed', 'passing', 'committing', 'creating PRs', 'verification', 'evidence'. It sits below anchor 5 because common variations like 'done', 'tests pass', 'build succeeds', or 'looks good' are not included.

4 / 5

Distinctiveness Conflict Risk

The trigger is anchored to a distinct moment — the act of claiming completion or committing/PR — which separates it from execution skills like a 'verify' or 'run tests' skill. It is not a 5 because it is a cross-cutting behavioral rule that fires near the end of nearly every task, creating minor overlap with any completion/verification-adjacent skill.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.