CtrlK
BlogDocsLog inGet started
Tessl Logo

verify-behavior

Owns a cheapest-first three-tier verification ladder — Tier 1 syntactic (grep / ast-grep / read), Tier 2 semantic-no-execution (typecheck / build / lint), Tier 3 execution (run the covering test, or a minimal synthesized repro) — and reports the result as an evidence receipt (confirms / contradicts / ambiguous / null). It never scores; `confidence(code)` owns the number. Two consumer shapes: claim-verification (read-only, feeds `confidence(code)`) and change-verification (post-apply green/red gate). Called by `verification-receipt.md` (pr-reviewer Tier 2/3), `bug-fix-verifier`, `feature-pr-verifier`, and the `aw-executor` Phase 4 checks loop. Use when a finding or a change needs executed proof, not just a plausible-sounding claim. Triggers on "verify this claim", "does this actually happen at runtime", "prove this behavior", "run this to confirm", "/verify-behavior".

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/quality/verify-behavior/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well structured as a thin index with a clearly sequenced V1–V5 workflow, explicit gates, and a checklist, but it is undermined by two issues: heavy repetition of the same invariants across the ladder tables, phase sections, Core Principles, anti-patterns, and Definition of Done, and — critically — every referenced rules/*.md file is missing from the bundle, so all operational detail (receipt format, discovery procedure, isolation commands) is unreachable. The result is an index pointing at content that does not ship.

Suggestions

Ship the referenced bundle files (rules/receipt.md, rules/ladder.md, rules/toolchain-discovery.md, rules/isolation-safety.md, rules/diagnostic-surface.md) — every link in the body is currently dead, so the progressive-disclosure promise is unfulfilled.

State each invariant once and cross-reference it: cheapest-first, isolation, and null-is-never-confirmation each appear in 4–5 sections (tables, phases, Core Principles, anti-patterns, Definition of Done); deduplicating these and the Required Reading table (which re-lists links already given inline) would cut the file substantially.

Inline the minimal executable essentials — the receipt format (fields plus the four verdict tokens) and a one-line worktree creation/cleanup command — so the skill remains actionable when the rule files are not loaded.

DimensionReasoningScore

Conciseness

There is no padding or explanation of concepts Claude already knows, but the same invariants are repeated many times: cheapest-first/no-Tier-3-escalation appears in the ladder table, the V2 section, Core Principle 1, and anti-pattern 1; the null-is-never-confirmation rule appears in the execute-not-score section, V4, Core Principle 6, anti-pattern 5, and the Definition of Done; and the 'Required Reading by Phase' table re-lists every rules/*.md link already given inline in V1–V5. This matches anchor 3 'mostly efficient but... could be tightened' rather than 4, since the redundancy goes beyond minor trimming for a self-described 'thin index'.

3 / 5

Actionability

The body contains concrete elements — real tool commands in the tier table ('tsc --noEmit', 'go build/go vet/staticcheck', 'cargo check/clippy', 'pyright/mypy'), explicit mode detection ('Parse the first token of $ARGUMENTS'), and concrete tier-selection heuristics ('A claim about symbol absence or a missing guard is usually a Tier 1 grep'). However the key executable details — the receipt format, the toolchain discovery procedure, and the worktree-isolation commands — are all deferred to rules files, and no complete executable procedure appears in the body. Matches anchor 3 'Some concrete guidance but incomplete... missing key details', not 4, because the specifics needed to actually produce a receipt are absent from this file.

3 / 5

Workflow Clarity

The V1–V5 phases are clearly sequenced with an explicit gate per phase in the workflow table ('Discovery order resolved', 'Cheapest tier that can decide the claim', 'Throwaway worktree... tracked files never modified'), an escalation feedback loop ('Walk the ladder Tier 1 → Tier 2 → Tier 3, stopping at the first tier that can decide the claim'), and a Definition of Done checklist ('Receipt returned with a verdict token; null/contradicting results dropped, never confirmed'). This matches anchor 5 'Clear sequence with explicit validation steps; feedback loops... checklists'. The destructive/batch cap does not apply — this skill is read-only or isolated by design.

5 / 5

Progressive Disclosure

The thin-index structure and signaling are well designed — 'Detailed rules live in `rules/*.md` and load on demand', a Required Reading by Phase table, and prominent per-phase links — but every referenced bundle file is missing: `rules/receipt.md`, `rules/ladder.md`, `rules/toolchain-discovery.md`, `rules/isolation-safety.md`, `rules/diagnostic-surface.md`, and `agents/shared/rules/verification-receipt.md` do not exist in the bundle. Scoring against the actual bundle structure, the navigation leads nowhere and all detailed content is absent. Not 4 ('most content is appropriately placed') since no referenced file resolves; not 2 since the index itself is clearly structured and references are prominently signaled, not buried or inlined.

3 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, highly specific description: it states concrete capabilities (three-tier ladder, tool examples, four-verdict receipt), an explicit use-when clause with natural trigger phrases, and clear boundaries against the adjacent confidence-scoring skill. The only minor weakness is that the internal wiring detail ('aw-executor Phase 4 checks loop' caller list) adds length without user-facing trigger value, and a few natural trigger synonyms are absent.

DimensionReasoningScore

Specificity

The description lists multiple concrete, specific actions with comprehensive coverage: 'Owns a cheapest-first three-tier verification ladder — Tier 1 syntactic (grep / ast-grep / read), Tier 2 semantic-no-execution (typecheck / build / lint), Tier 3 execution (run the covering test, or a minimal synthesized repro) — and reports the result as an evidence receipt (confirms / contradicts / ambiguous / null)'. Concrete tool names, tier structure, and verdict tokens are all explicitly stated; it also defines both consumer shapes ('claim-verification (read-only...)' and 'change-verification (post-apply green/red gate)'). This matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage' rather than 4, which would require minor coverage gaps — none are evident.

5 / 5

Completeness

Both what and when are explicitly answered: what — 'Owns a cheapest-first three-tier verification ladder... reports the result as an evidence receipt'; when — 'Use when a finding or a change needs executed proof, not just a plausible-sounding claim' followed by concrete trigger phrases. This matches the anchor 'Clearly and explicitly answers both what AND when with concrete trigger phrases' — the 'when' clause is explicit, not implied, so 4's 'when could be more explicit' does not apply.

5 / 5

Trigger Term Quality

Explicit natural trigger phrases are present: 'Triggers on "verify this claim", "does this actually happen at runtime", "prove this behavior", "run this to confirm", "/verify-behavior"', plus the 'Use when a finding or a change needs executed proof' clause. These are phrases a user would naturally say, with synonyms covered (verify/prove/confirm). Not a 5 because a few common natural variations are missing (e.g. 'test this', 'reproduce the bug', 'confirm the fix'); the set is good but not exhaustively comprehensive.

4 / 5

Distinctiveness Conflict Risk

The description carves out a clear niche — executed verification with a receipt — and explicitly fences itself off from the adjacent skill: 'It never scores; `confidence(code)` owns the number'. Trigger phrases like 'prove this behavior' and 'does this actually happen at runtime' are distinctive of this niche. Minimal conflict risk; matches anchor 5 'Clear niche with distinct triggers'. Not 4, since the overlap risk with generic test-running skills is directly addressed by the never-scores boundary and named consumer list.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 18 missing, 1 suspicious

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.