CtrlK
BlogDocsLog inGet started
Tessl Logo

prove-it-works

Prove a task against the real artifact before calling it done. Use during factory-work, before ready_for_review, and whenever a check only shows that code compiles or a file changed.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary short, single-purpose skill: lean prose, an executable evidence-manifest specification, explicit validation criteria, and a clear failure heuristic with no padding or dangling references.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence — "Verify the real artifact. Do not infer from proxies, self-reports, or 'it compiles.'" — with no explanation of known concepts; the only example (the manifest JSON) teaches a project-specific format Claude could not already know. Every token earns its place.

5 / 5

Actionability

Guidance is fully concrete: an exact output path ('.factory/evidence/<ticket>/<task>/manifest.json'), a copy-paste-ready manifest example with command, cwd, exitCode, and artifact, a hard acceptance rule ("A criterion without a command and a zero exit code is not proof"), and a specific UI variant ("the Surf screenshot and the read-back of the rendered page, taken against the worktree's own server").

5 / 5

Workflow Clarity

This is a simple, single-purpose skill and the single action is unambiguous: check the real artifact, then record evidence. It includes an explicit validation standard (zero exit code) and an error-recovery heuristic ("When a check fails, suspect the observation method before suspecting the product"), so the simple-skill exception applies cleanly.

5 / 5

Progressive Disclosure

The skill is 37 lines, needs no external references (none exist and none are referenced), and is well organized as principle → verification checklist → evidence format → commit rule. Per the under-50-line guideline, this earns the top score without section headers or bundle files.

5 / 5

Total

20

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, well-formed description with an explicit 'Use during…' trigger clause and concrete completion criterion. Its main limitation is that it names a single action and relies on environment-specific jargon rather than enumerating capabilities or broader natural synonyms.

Suggestions

Name the concrete verification actions in the description (e.g., 'run the feature, read the actual value, record the command and exit code as evidence') so the capability list matches the body's specificity.

Add one or two widely natural trigger synonyms such as 'verify', 'test', or 'evidence' so the skill is discoverable by users who don't know the terms 'factory-work' or 'ready_for_review'.

DimensionReasoningScore

Specificity

The description states one crisp concrete action — "Prove a task against the real artifact before calling it done" — and scopes it with trigger conditions rather than listing distinct capabilities. It matches the anchor 'names domain and 1-2 concrete actions, but not comprehensive'; it is not 4 because it does not enumerate several specific actions.

3 / 5

Completeness

It explicitly answers what ("Prove a task against the real artifact before calling it done") and when ("Use during factory-work, before ready_for_review, and whenever a check only shows that code compiles or a file changed") with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Triggers like "whenever a check only shows that code compiles or a file changed" and "before ready_for_review" are natural phrases in this workflow, giving good keyword coverage. It falls short of 5 because common synonyms such as 'verify', 'test', or 'evidence' are absent and coverage leans on project-specific jargon ('factory-work').

4 / 5

Distinctiveness Conflict Risk

The evidence-manifest niche with triggers tied to 'factory-work' and 'ready_for_review' is mostly distinct, but 'prove it works' overlaps slightly with generic run/verify-type skills, so minor overlap risk keeps it at 4 rather than 5.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
geut/factory-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.