Owns a cheapest-first three-tier verification ladder — Tier 1 syntactic (grep / ast-grep / read), Tier 2 semantic-no-execution (typecheck / build / lint), Tier 3 execution (run the covering test, or a minimal synthesized repro) — and reports the result as an evidence receipt (confirms / contradicts / ambiguous / null). It never scores; `confidence(code)` owns the number. Two consumer shapes: claim-verification (read-only, feeds `confidence(code)`) and change-verification (post-apply green/red gate). Called by `verification-receipt.md` (pr-reviewer Tier 2/3), `bug-fix-verifier`, `feature-pr-verifier`, and the `aw-executor` Phase 4 checks loop. Use when a finding or a change needs executed proof, not just a plausible-sounding claim. Triggers on "verify this claim", "does this actually happen at runtime", "prove this behavior", "run this to confirm", "/verify-behavior".
64
80%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./skills/quality/verify-behavior/SKILL.mdGiven a behavioral claim about code, or a change that was just applied, decide the cheapest way to get executed proof, run it in isolation, and report the raw result as a receipt.
This skill is the execution engine six-plus call sites in this repo used to hand-roll independently: "detect the toolchain, run something, read pass or fail." It replaces the ad hoc version in each of those with one shared ladder.
This
SKILL.mdis a thin index. Detailed rules live inrules/*.mdand load on demand.
This skill does not score — it never assigns a confidence score and never grades pass/fail against an intent.
It runs a command, captures the raw output, and classifies the result against the claim itself as confirms / contradicts / ambiguous / null.
confidence(code) owns the number — this skill supplies sharper evidence to that gate, it does not replace it.bug-fix-verifier's FAIL_TO_PASS, the aw-executor Phase 4 expect comparison) stays with the caller — this skill supplies the run-and-observe mechanic underneath that grading, not the grading itself.See rules/receipt.md for the full contract, including the hard invariant that a null or non-reproducing result drops or contradicts a finding and is never confirmation.
| Shape | Question it answers | Consumers |
|---|---|---|
| Claim-verification | "Is this specific behavioral assertion true?" — read-only, feeds confidence(code) as Evidence | agents/shared/rules/verification-receipt.md (pr-reviewer Tier 2/3) |
| Change-verification | "Did this applied change produce the expected green/red result?" — a post-apply gate | bug-fix-verifier, feature-pr-verifier, aw-executor Phase 4 checks loop |
Both shapes share the same core: toolchain discovery, isolated execution, and the receipt format.
Only the output framing differs — see rules/receipt.md.
Parse the first token of $ARGUMENTS.
| Mode | Default | Trigger | What it does |
|---|---|---|---|
claim | yes | No mode token, or claim | Verify one behavioral assertion. Returns a single receipt. Read-only. |
change | First token change | Verify a just-applied change against an expected outcome. Returns a green/red gate result with the same receipt shape underneath. |
claim (claim mode) — the behavioral assertion in prose, e.g. "validateAuth throws on an invalid token."target — the file(s) or symbol the claim or change concerns.expected (change mode) — the expected post-change outcome (e.g. an expect: string from checks.yaml, or "the repro now passes").review_relation — "self" | "cross" | "untrusted" (default "self" for a caller's own branch).
Governs the Tier 3 trust split — see rules/isolation-safety.md.caller — the invoking agent or skill, for logging only.| Tier | Name | Cost | Example tools |
|---|---|---|---|
| Tier 1 | Syntactic | Lowest | grep, ast-grep, Read |
| Tier 2 | Semantic-no-execution | Low | tsc --noEmit, go build/go vet/staticcheck, cargo check/clippy, pyright/mypy |
| Tier 3 | Execution | Highest | Run the covering test, or synthesize and run a minimal repro |
Stop at the cheapest tier that can decide the claim — do not escalate to Tier 3 when Tier 1 or Tier 2 already confirms or contradicts it.
Full per-language mapping in rules/ladder.md.
| Phase | Name | Rule file | Gate |
|---|---|---|---|
| V1 | Toolchain discovery | rules/toolchain-discovery.md | Discovery order resolved; never assume a global install |
| V2 | Tier selection | rules/ladder.md | Cheapest tier that can decide the claim, per the per-language adapter table |
| V3 | Isolated execution | rules/isolation-safety.md | Throwaway worktree (Tier 3), tracked files never modified, scratch deleted, relation-keyed trust split honored |
| V4 | Receipt | rules/receipt.md | confirms/contradicts/ambiguous/null; null-is-never-confirmation invariant |
| V5 | Report | this file + rules/receipt.md | Claim mode returns the receipt; change mode returns the receipt plus a green/red verdict |
Resolve what to run before deciding how to run it: checks.yaml first, then the argent-environment-inspector detection pattern, then manifest scripts.
Never assume a tool is globally installed.
See rules/toolchain-discovery.md.
Walk the ladder Tier 1 → Tier 2 → Tier 3, stopping at the first tier that can decide the claim.
A claim about symbol absence or a missing guard is usually a Tier 1 grep.
A claim about a type contract is usually a Tier 2 typecheck.
A claim about runtime return value, thrown error, or side-effect ordering needs Tier 3.
See rules/ladder.md for the full per-language table.
Tier 3 runs in a throwaway worktree, never touches tracked files, deletes its scratch harness after, defaults to no network, and never pipes a remote script into a shell.
Tier 3 is default-on only for the caller's own code (self relation); cross/untrusted callers need an explicit sandbox opt-in.
See rules/isolation-safety.md.
Every run — regardless of tier or mode — produces a receipt: the raw command, its raw output, and one of four verdict tokens.
A null or empty result is dropped or contradicts; it is never read as confirmation.
See rules/receipt.md.
confidence(code) Evidence).expected outcome.
The caller keeps its own grading semantics on top (e.g. checks.yaml's expect: comparison, FAIL_TO_PASS).Load on demand — do not preload.
| Phase | Files |
|---|---|
| V1 | rules/toolchain-discovery.md |
| V2 | rules/ladder.md |
| V3 | rules/isolation-safety.md |
| V4, V5 | rules/receipt.md |
| wiring | agents/shared/rules/verification-receipt.md — how pr-reviewer calls this skill |
| diagnose | rules/diagnostic-surface.md |
rules/toolchain-discovery.md).curl | sh (rules/isolation-safety.md).grep could decide.tsc/go/cargo/pyright is on PATH without checking the project's actual toolchain.review_relation.39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.