CtrlK
BlogDocsLog inGet started
Tessl Logo

verify

AI DevKit · Enforce evidence-based completion claims — require fresh command output before reporting success. Use when completing any task, fixing a bug, finishing a phase, running tests, building, deploying, or making any "it works" claim.

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable instruction skill with explicit validation checkpoints, feedback loops, and concrete commands. It is token-efficient and well-structured with no unnecessary reference indirection.

DimensionReasoningScore

Conciseness

Lean body with no padding or re-explanation of concepts Claude already knows; every section (Hard Rules, Gate, Patterns, Regression, Red Flags) earns its place.

3 / 3

Actionability

Concrete, copy-paste-ready guidance including the 'npx ai-devkit@latest memory store ...' command, a precise 5-step gate, and a 6-step regression procedure with explicit pass/fail criteria.

3 / 3

Workflow Clarity

The 5-step gate and regression verification are clearly sequenced with explicit validation checkpoints and feedback loops ('If any step fails, stop. Fix and restart from step 1'; 'If step 4 passes, the test is wrong').

3 / 3

Progressive Disclosure

Single self-contained file with well-organized, clearly headed sections and no nested references; appropriate for a focused skill with no bundle files.

3 / 3

Total

12

/

12

Passed

Description

92%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concrete description with explicit 'Use when' triggers and natural language terms. Its only weakness is trigger breadth: 'completing any task' risks firing alongside many other skills.

Suggestions

Narrow the 'Use when' triggers from 'completing any task' to verification-specific moments (e.g., 'before reporting task/bug/test/build success') to reduce overlap with other skills.

Consider adding a negative boundary (e.g., 'Not for routine status updates') to improve distinctiveness.

DimensionReasoningScore

Specificity

Names a concrete core action ('require fresh command output before reporting success') and enumerates specific scenarios (fixing a bug, running tests, building, deploying), matching the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

Explicitly answers both what ('Enforce evidence-based completion claims — require fresh command output') and when via a clear 'Use when...' clause enumerating triggers.

3 / 3

Trigger Term Quality

Natural user-language triggers like 'fixing a bug', 'running tests', 'building', 'deploying', and 'it works' give good coverage of terms a user would actually say.

3 / 3

Distinctiveness Conflict Risk

The verification niche is distinct, but the trigger 'completing any task, fixing a bug, finishing a phase, running tests, building, deploying' is extremely broad and would overlap with nearly every task-oriented skill.

2 / 3

Total

11

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
codeaholicguy/ai-devkit
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.