CtrlK
BlogDocsLog inGet started
Tessl Logo

verifying-changes

Concrete, per-area proof that a change actually works before reporting it fixed, done, or "should work now" — which dev server, test command, or invocation proves a template UI change, an action, a migration, a guard, or a core/package change. Use before every wrap-up, and before stopping mid-task to ask permission instead of continuing.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplar of actionable, repo-specific guidance: a per-area proof table with exact commands, explicit validation and feedback loops, and crisp escalation rules. Its only soft spots are mild token overhead from the motivational quote block and a forensics section that is tangential to the core verification flow.

DimensionReasoningScore

Conciseness

The body is lean and imperative with no re-teaching of concepts Claude already knows; every section drives a decision or command. The "Real failures this replaces" quote block (~13 lines of rationale) is the one padding instance that could be trimmed without losing executable guidance, keeping it just below anchor 5.

4 / 5

Actionability

Every change area maps to a copy-paste-ready command: `pnpm --filter <app> dev`, `cd templates/<app> && pnpm action <name> --key value`, `pnpm guard:<name>`, `node --experimental-strip-types --test scripts/ci-change-scope.test.ts`, and `pnpm run prep`, plus the `db-query` read-back to confirm writes landed — specific examples cover the common cases exactly as anchor 5 requires.

5 / 5

Workflow Clarity

The workflow is clearly sequenced (pick the smallest proof by area, run it, show the proof in the reply, state the gap if you cannot run it) with explicit validation steps (db read-back after writes, console + network checks), feedback loops ("A failing check is a reason to keep working", re-run against the latest edit rather than memory), and checklists (the "Do NOT report done" list and per-area table). On anti-drift re-read, no validation gaps remain to justify anchor 4.

5 / 5

Progressive Disclosure

There is no bundle (no references/, scripts/, or assets/) and the one cross-skill reference ("see `adding-tests-and-ci`") is clearly signaled, with well-organized sections throughout. At ~110 lines, the orthogonal "Production forensics" section and the quotes section could live in separate reference files, which keeps it at anchor 4 rather than 5.

4 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what it does and when to use it, with concrete per-area vocabulary and low conflict risk. Its main gap is trigger phrasing: it leans on internal process timing ("before every wrap-up") rather than the natural words a user would say ("verify", "end to end", "is it fixed").

Suggestions

Add natural user-said trigger terms to the description — e.g., "verify", "confirm", "end to end", "is it actually fixed" — so the skill surfaces when the user asks for proof, not just at wrap-up time.

Rephrase the when-clause to include a concrete user-facing trigger (e.g., "Use before every wrap-up and whenever the user asks whether a change is fixed or tested") in addition to process timing.

Mention the remaining proof areas from the body (CI/workflow changes, cross-cutting changes, docs-only) so the what-clause's coverage is comprehensive.

DimensionReasoningScore

Specificity

Names concrete proof mechanisms ("which dev server, test command, or invocation") and enumerates specific change areas ("a template UI change, an action, a migration, a guard, or a core/package change"). It lists several specific actions but omits areas the body covers (CI jobs, deployed behavior, docs), so it falls just short of comprehensive anchor 5.

4 / 5

Completeness

Explicitly answers both what ("Concrete, per-area proof that a change actually works") and when ("Use before every wrap-up, and before stopping mid-task"), so the missing-when cap of 3 does not apply. The when-clause is framed as internal process timing rather than the concrete user-mentioned trigger phrases of anchor 5.

4 / 5

Trigger Term Quality

Includes natural phrases users say in this situation ("works", "fixed, done, or 'should work now'", "test command"), giving good keyword coverage. Common synonyms like "verify", "confirm", or "end to end" are missing, which keeps it below anchor 5's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — proving changes work before reporting done — with repo-specific trigger vocabulary (template UI, actions, migrations, guards, core/package) that no generic testing or document skill would match, so conflict risk is minimal.

5 / 5

Total

17

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

12

/

16

Passed

Repository
BuilderIO/agent-native
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.