CtrlK
BlogDocsLog inGet started
Tessl Logo

verify-first

Ground-truth every technical claim — and every requirement-coverage claim — in a plan or design against the source of truth (run the code, hit the API with curl, probe the live browser DOM/store, build a proof-of-concept; for requirements, diff the plan section-by-section against the original design doc + ticket) before it can be relied on. Reading docs, blogs, training knowledge, or trusting that the plan captured the requirements is NOT verification, and neither is a probe that could not have failed (an HTTP 200, a bound port, a detector that reported nothing because it never ran). Use when stress-testing assumptions, before locking an architectural decision, or when the user says "verify first", "verify with practical tests", "prove it", or "test, don't assume".

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is verify-first in AndreJorgeLopes/devflow

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers exceptional workflow discipline and actionable, feedback-looped guidance — the [V]/[B]/[R] ledger system with blocking semantics is a model of concrete instruction. Its weaknesses are repetition (lessons stated twice, once as prose and again in the anti-patterns table) and a monolithic structure where the war stories could be split into reference files to keep the core checklist lean.

Suggestions

Merge the three prose 'lessons' with the anti-patterns table — keep one authoritative version (the table is the more scannable form) and cut the duplicated narrative, saving roughly 30 lines of always-loaded context.

Move the two case studies ("The override that started this skill" and the false-positive app story) into a references/ file (e.g. references/case-studies.md), keeping a one-line distilled lesson in SKILL.md with a clearly signaled link.

Tighten the 'Liveness is not correctness' table and 'Detectors fail silent-clean' section into a single compact checklist of probe-validity rules; their current two-section form restates the same idea twice.

DimensionReasoningScore

Conciseness

The body is genuinely non-obvious in places ("A green detector is evidence only after you have watched it go RED on a known-bad input"), but it is noticeably padded: the two war-story sections ("The override that started this skill", "The inverse failure") run long, and the lessons taught there are then restated almost verbatim in the 12-row anti-patterns table ("One negative probe ≠ absence", "200 is liveness, not correctness", "Solved problem = untested edge-case claim"). This fits 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than anchor 4, since whole sections are duplicated.

3 / 5

Actionability

For an instruction-only skill the guidance is fully actionable: copy-paste-ready probes ("curl -s $URL | jq .nextToken", the CometRelayEnvironment browser snippet), a complete fill-in verification-ledger template with worked example rows, a library-inventory template, an explicit probe-selection heuristic ("if the thing I am checking were broken, would this probe look any different?"), and an eval-mode guard with exact behavioral rules. Per the code-vs-instruction scoring note, absence of code is not penalized when guidance is this concrete.

5 / 5

Workflow Clarity

The 6-step workflow is clearly sequenced, turned into a TodoWrite checklist with per-claim todos, and has explicit validation checkpoints and feedback loops: [B] tests are BLOCKING, a negative probe triggers a second differently-shaped probe, evidence must be captured before [V], and Step 6 is an explicit gate ("A decision may NOT be locked… while any claim it depends on is unclassified or still [B]-unrun") with [R] items carried to a closure gate. This matches the anchor-5 pattern of clear sequence with explicit validation and error-recovery loops.

5 / 5

Progressive Disclosure

The skill is a single ~170-line file with no bundle directories and no external references; the section headers are sensible but content that would sit better in a separate reference file — the two multi-paragraph case studies and the liveness-vs-correctness table — is inlined in the always-loaded SKILL.md. This fits 'some structure but could be better organized; content that should be separate is inline' rather than anchor 4, which presumes appropriate placement across files.

3 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete actions, an explicit 'Use when' clause with verbatim user trigger phrases, and sharp boundary-setting (what does NOT count as verification, including probes that could not have failed). Its only weaknesses are length — it is a single dense sentence-pair that could be tightened — and a few missing natural trigger synonyms.

DimensionReasoningScore

Specificity

The description lists multiple concrete, specific actions — "run the code, hit the API with curl, probe the live browser DOM/store, build a proof-of-concept; for requirements, diff the plan section-by-section against the original design doc + ticket" — comprehensively covering both technical and requirement-coverage claims. This matches the anchor 'lists multiple specific concrete actions; comprehensive coverage' rather than anchor 4, since coverage has no meaningful gaps.

5 / 5

Completeness

It explicitly answers both questions: the "what" is grounding every technical and requirement-coverage claim against the source of truth, and the "when" is an explicit "Use when stress-testing assumptions, before locking an architectural decision, or when the user says…" clause with concrete trigger phrases. This is a clear match for the anchor-5 exemplar.

5 / 5

Trigger Term Quality

Strong natural trigger phrases are quoted verbatim — "verify first", "verify with practical tests", "prove it", "test, don't assume", plus "stress-testing assumptions" — so users would naturally say these. It falls short of anchor 5 only because a few common variants (e.g. "double-check this", "is this actually true", "validate this assumption") are absent.

4 / 5

Distinctiveness Conflict Risk

The verification/ground-truthing niche and quoted triggers ("prove it", "verify first") are distinctive, but phrases like "prove it" and "test, don't assume" carry minor overlap risk with general testing or code-review skills. It sits between anchors 4 and 5; 'mostly distinct; minor overlap risk with closely related skills' is the closer fit.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
AndreJorgeLopes/devflow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.