CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-verification-gate

Use when about to declare work complete, fixed, passing, or done

51

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-verification-gate/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, disciplined behavioral gate: a crisp step sequence with hard validation checkpoints, concrete evidence requirements per claim type, and a correct/incorrect example pair. Its weaknesses are redundancy (four sections restate the same rule) and a dangling cross-reference section pointing at files outside the bundle.

Suggestions

Merge the Rationalization Table and the Red Flags table — both map the same excuses/thoughts to "run fresh verification" — and fold "When to Apply" into the Gate section to cut roughly a third of the body.

Drop or prune the Multi-Provider Context section (or move its bash checks into the Evidence table); it is niche padding for most invocations.

Remove or verify the 'Integration with Other Skills' cross-references — none of the six named files exist in this bundle, so the section is unresolvable for the model.

DimensionReasoningScore

Conciseness

The body is terse and never explains concepts Claude already knows, but the message "run fresh verification before claiming success" is restated across four overlapping sections — the Rationalization Table ("I ran the tests earlier this session… Earlier is not fresh"), the Red Flags table ("Should work now → Run the verification"), "When to Apply", and The Gate — plus a niche Multi-Provider section that could be tightened. This is "mostly efficient but could be tightened" rather than minor trimming.

3 / 5

Actionability

Concrete, executable guidance dominates: the five-step gate (IDENTIFY/RUN/READ/VERIFY/ONLY THEN), copy-paste bash checks ("ls -la ~/.claude-octopus/results/*-synthesis-*.md | tail -1", "wc -l …"), an evidence table mapping each claim to its required output, and a worked red-green sequence. It misses anchor 5 only because the core directive "execute the full command" stays generic (no per-context command guidance beyond the multi-provider case).

4 / 5

Workflow Clarity

The Gate gives a clear, numbered five-step sequence with an explicit validation checkpoint ("VERIFY — Does output actually confirm the claim?"), a hard skip rule ("Skip any step = the claim is unverified"), a failure-path example (red → green → revert-to-red → green proves the test isn't a false positive), and per-phase checkpoints for the orchestrate.sh workflow. This matches the anchor for clear sequence with explicit validation and feedback loops.

5 / 5

Progressive Disclosure

The body is well-organized with clear section headers, appropriately sized for a single-purpose gate skill, and no bundle files exist to reference — the whole content reasonably lives inline. It misses anchor 5 because the "Integration with Other Skills" section lists six sibling files (flow-develop.md, skill-tdd.md, etc.) that are not part of this bundle, leaving dangling references the model cannot navigate.

4 / 5

Total

16

/

20

Passed

Description

40%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a well-formed trigger clause with natural keywords, but it is only a trigger — the skill's actual behavior (run fresh verification, read output, gate claims on evidence) is entirely unstated. This makes it incomplete and overly broad for triggering.

Suggestions

Add a 'what' clause naming the skill's concrete actions, e.g. "Runs the verification command fresh, reads full output, and gates completion claims on evidence."

Broaden trigger coverage with common synonyms such as 'verified', 'tests pass', 'all green', 'before committing or shipping'.

Sharpen distinctiveness by scoping the trigger, e.g. "before committing, marking a task done, or reporting results" to reduce overlap with general test/verification skills.

DimensionReasoningScore

Specificity

The description is only "Use when about to declare work complete, fixed, passing, or done" — it implicitly names the domain (gating completion claims) but lists zero concrete actions or capabilities. It sits above anchor 1 ("pure abstract language" like "Helps with documents") because the trigger context is identifiable, but below anchor 3, which requires naming 1-2 concrete actions.

2 / 5

Completeness

Only the "when" is present ("Use when about to declare work complete, fixed, passing, or done"); there is no "what" — the description never states what the skill does (run verification commands and check evidence before making claims). This matches anchor 2 ("only 'when' is present without 'what'", e.g. "Use when working with documents").

2 / 5

Trigger Term Quality

"complete, fixed, passing, or done" are natural words a user or Claude would actually say when this skill is needed, giving good keyword coverage. It falls short of anchor 5 because common variations like "verified", "tests pass", "ship", or "before committing" are absent.

4 / 5

Distinctiveness Conflict Risk

"About to declare work complete, fixed, passing, or done" applies to essentially every coding session's end state, so it would fire alongside any test-running, verification, or code-review skill — high overlap risk, matching anchor 2 ("Very broad; high overlap risk"). It is not anchor 1 ("conflicts with virtually any skill"), because the completion-claim context does carve out a niche.

2 / 5

Total

10

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.