CtrlK
BlogDocsLog inGet started
Tessl Logo

jev-assert

Verifies a UI expectation semantically by asking TypeSafe's Jev model whether a user-observable outcome holds in a page's captured TEXT state (accessibility tree / page text — never a screenshot), then maps Jev's typed answer and probability onto this repo's verify-behavior receipt vocabulary (confirms, contradicts, ambiguous, null, unobtainable). Driver-agnostic: consumes text state from either the Playwright (aw-tester) or the claude-in-chrome (aw-tester-chrome) runner, so the Chrome-vs-Playwright choice is orthogonal to it. Use it for a semantic UI assertion where an exact locator or string match is brittle — "did the user see a success state", "is this the right screen". Delegates the API contract to the typesafe@typesafe-ai skill (TYPESAFE_API_KEY). Triggers on "assert semantically", "does the page show", "verify this outcome with jev", "semantic UI assertion", "jev-assert", "/jev-assert".

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-written as an index: a tight receipt example, an executable call path, a gated three-step workflow, and honest failure semantics. Its decisive weakness is packaging, not prose — the rules/*.md files it designates as the substance of the skill are absent from the bundle, so the progressive-disclosure structure it advertises cannot actually be navigated.

Suggestions

Ship the referenced rule files (rules/provenance.md, rules/state-extraction.md, rules/receipt-mapping.md) in the bundle, or inline their essential content (threshold defaults/bands, provenance rejection criteria, state-extraction requirements) into SKILL.md — currently every workflow gate points to a file that does not exist.

Consolidate the repeated text-only rule (stated in 'Why Jev, and why text-only', the Step 1 gate, Core Principle #3, and Anti-patterns) and the repeated unobtainable/never-guess rule into one authoritative statement each to save tokens.

Add per-gate recovery guidance — e.g., when Step 0 rejects an expectation as by-construction, instruct the executor to rephrase it as a user-observable outcome and re-run — to turn the gates into a feedback loop rather than dead ends.

DimensionReasoningScore

Conciseness

The body is a genuinely lean "thin index" — the invocation contract and command example carry real contract information with no tutorial padding about concepts Claude already knows. It is not a 5 because the text-only rule is restated roughly four times ("Jev evaluates text, not images", the Step 1 gate "no screenshot", Core Principle #3 "Jev never sees a screenshot", and Anti-pattern #1 "Feeding Jev a screenshot"), and the fail-honest rule is similarly repeated — trimmable repetition, but only minor.

4 / 5

Actionability

The concrete call path is copy-paste ready — `node ${CLAUDE_SKILL_DIR}/scripts/jev-call.mjs --expectation "the user sees an order-confirmation number" --state-file <captured-state.txt> --target <page-or-spec-id>` — and `--self-test` gives an offline verification command, with the script actually present in the bundle. It is not a 5 because the decision rules the flags depend on (threshold bands, provenance rejection criteria, state-extraction requirements) are delegated to rules files that are not in the bundle, leaving gaps an executor cannot close.

4 / 5

Workflow Clarity

The workflow is a clearly sequenced three-step table with an explicit Gate column per step ("Expectation is a user-observable OUTCOME, not a restatement of the state text"; "Page TEXT state captured... no screenshot"; "mapped to exactly one receipt verdict") plus a defined failure path ("When TYPESAFE_API_KEY is unset or the API is unreachable... the script emits unobtainable, never a confirms / contradicts guess"). It is not a 5 because there is no feedback loop: what the executor should do when a gate fails (e.g., rephrase a rejected expectation) is never stated.

4 / 5

Progressive Disclosure

The body claims "This SKILL.md is a thin index. Detailed rules live in rules/*.md and load on demand" and links rules/provenance.md, rules/state-extraction.md, rules/receipt-mapping.md, scripts/validate-skill.mjs, evals/evals.json, and evals/triggers.jsonl — but the actual bundle contains only scripts/jev-call.mjs; none of the referenced files exist. Scored against the real bundle structure, the reference architecture is broken: every gate-owning rule file is a dangling pointer, recreating the "See advanced.md for details" anti-pattern. It is not a 3 because anchor 3's flaw (references present but not clearly signaled) is far milder than references that resolve to nothing, and not a 1 because the links are one level deep and the body does inline enough contract detail to run the happy path.

2 / 5

Total

14

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete actions, an explicit use-when clause with natural trigger phrases, and a clearly delineated niche. The only soft spot is slight overlap risk with the repo's other verification skills, which share the receipt vocabulary it emits.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Verifies a UI expectation semantically by asking TypeSafe's Jev model", "maps Jev's typed answer and probability onto this repo's verify-behavior receipt vocabulary (confirms, contradicts, ambiguous, null, unobtainable)", and "consumes text state from either the Playwright (aw-tester) or the claude-in-chrome (aw-tester-chrome) runner" — with comprehensive coverage of the skill's behavior. It is not a 4 because there are no coverage gaps: input contract, output vocabulary, drivers, and API-key delegation are all named.

5 / 5

Completeness

Both questions are answered explicitly: the 'what' is the verify-and-map pipeline stated in the first sentence, and the 'when' is explicit — "Use it for a semantic UI assertion where an exact locator or string match is brittle" followed by a concrete trigger-phrase list. Not a 4 because the 'when' clause is already fully explicit with concrete trigger phrases, leaving nothing merely implied.

5 / 5

Trigger Term Quality

Trigger coverage is comprehensive and natural: "Triggers on 'assert semantically', 'does the page show', 'verify this outcome with jev', 'semantic UI assertion', 'jev-assert', '/jev-assert'", plus user-voice paraphrases like "did the user see a success state" and "is this the right screen". It includes both the natural phrasing and the exact command name, matching the anchor-5 example's synonym coverage.

5 / 5

Distinctiveness Conflict Risk

It carves a clear niche (semantic UI assertion via Jev over text state) with distinctive triggers like "assert semantically" and "verify this outcome with jev", but it operates in the same verification territory as the repo's named sibling skills (verify-behavior, ui-verify) and shares their receipt vocabulary, leaving minor overlap risk with closely related skills. It is not a 5 because a query like "verify this outcome" could plausibly be intended for verify-behavior instead.

4 / 5

Total

19

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 5 missing, 6 suspicious

Warning

referenced_paths_exist

Referenced path issues: 1 missing, 1 deeper-than-1-level

Warning

Total

12

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.