CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-verification-gate

Use when about to declare work complete, fixed, passing, or done

53

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-verification-gate/SKILL.md

The canonical home for this skill is skill-verify in nyldn/claude-octopus

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strongly actionable behavioral skill: a clear sequenced verification workflow with explicit validation checkpoints, evidence requirements, feedback loops, and concrete examples. Its weaknesses are a redundant pair of anti-rationalization sections and cross-references to sibling files that don't exist in the bundle.

DimensionReasoningScore

Conciseness

The body is mostly efficient — tables, terse imperatives, no concept explanations Claude doesn't need — but the 'Rationalization Table' and 'Red Flags — STOP and Verify' sections repeat the same items ('I'm confident' and the subagent-excuse rows appear in both), which could be merged. Only one padded theme keeps it above anchor 2.

3 / 5

Actionability

Concrete, executable guidance throughout: the 5-step gate (IDENTIFY/RUN/READ/VERIFY/ONLY THEN), the claim-to-evidence table, real bash commands for the synthesis-file checks, and correct/incorrect output examples. The general-case verification command is necessarily project-specific and left implicit, a minor gap that keeps it below the copy-paste-ready bar of a 5.

4 / 5

Workflow Clarity

The Gate is a clearly sequenced 5-step procedure that is itself a validation sequence (READ full output, check exit code, count failures, VERIFY the output confirms the claim). The red-green regression example is an explicit feedback loop (revert fix → fail → restore), and the Evidence and When-to-Apply tables act as checklists.

5 / 5

Progressive Disclosure

A single-file skill (~130 lines) with clear, well-ordered section headers and no nested references — content is appropriately placed inline for its size. Minor gap: the host note and 'Integration with Other Skills' reference files (skills/blocks/codex-host-adapter.md, flow-develop.md, skill-tdd.md) that are not present in the bundle.

4 / 5

Total

16

/

20

Passed

Description

43%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a pure trigger clause: it says precisely when to invoke the skill but never says what the skill does, forcing the reader to infer the verification-gate capability. Trigger phrasing is natural and synonym-rich, but breadth creates moderate overlap risk with other verification skills.

Suggestions

Add an explicit 'what' clause naming the skill's actions, e.g. 'Requires fresh command evidence before any completion claim. Use when about to declare work complete, fixed, passing, or done.'

Include verification-adjacent trigger terms such as 'verify', 'tests pass', or 'before committing or reporting results' to broaden natural keyword coverage.

State the evidence rule (e.g. 'claims must cite command output from the current turn') in the description to sharpen distinctiveness from generic testing or code-review skills.

DimensionReasoningScore

Specificity

The description names the trigger situation concretely ('declare work complete, fixed, passing, or done') but states no actions the skill performs — there is no verb describing what it does (verify, gate, run). It is not pure abstract language (ruling out 1), but it lists none of the skill's concrete actions (ruling out 3).

2 / 5

Completeness

Only a 'when' clause is present ('Use when about to declare work complete, fixed, passing, or done'); the 'what' — that the skill gates completion claims behind fresh verification evidence — is entirely absent and must be inferred. This matches anchor 2 ('only when is present without what') exactly; anchor 3 requires a clear 'what', and none is explicitly stated.

2 / 5

Trigger Term Quality

'complete, fixed, passing, or done' gives good natural coverage of completion-claim synonyms an assistant or user would actually say. Missing verification-adjacent terms like 'verify', 'tests pass', or 'before committing', so it falls short of the comprehensive synonym coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

The trigger moment is somewhat specific, but 'about to declare work complete' fires at the end of nearly every task and overlaps with verification-type skills (e.g., a project's own verify/test skill). It could still collide with similar skills, matching anchor 3 rather than the minor-overlap profile of anchor 4.

3 / 5

Total

11

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.