Verifies a UI expectation semantically by asking TypeSafe's Jev model whether a user-observable outcome holds in a page's captured TEXT state (accessibility tree / page text — never a screenshot), then maps Jev's typed answer and probability onto this repo's verify-behavior receipt vocabulary (confirms, contradicts, ambiguous, null, unobtainable). Driver-agnostic: consumes text state from either the Playwright (aw-tester) or the claude-in-chrome (aw-tester-chrome) runner, so the Chrome-vs-Playwright choice is orthogonal to it. Use it for a semantic UI assertion where an exact locator or string match is brittle — "did the user see a success state", "is this the right screen". Delegates the API contract to the typesafe@typesafe-ai skill (TYPESAFE_API_KEY). Triggers on "assert semantically", "does the page show", "verify this outcome with jev", "semantic UI assertion", "jev-assert", "/jev-assert".
66
83%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Low
Low-risk findings worth noting
Turns a natural-language, user-observable UI expectation into an executed
semantic verdict.
It reads a page's captured text state, asks TypeSafe's Jev model one
narrow judgment, and emits a verify-behavior receipt the rest of this repo
already consumes.
It never drives a browser, never applies a fix, and never writes code — it is a
read-only verification primitive, the semantic sibling of
verify-behavior.
This
SKILL.mdis a thin index. Detailed rules live inrules/*.mdand load on demand.
The terminal deliverable is one receipt line, byte-compatible with the
canonical vocabulary in
verification-receipt.md:
[receipt] tier: 3 | tool: jev | target: <page/spec>
[receipt] question: Does the page state show that <expectation>?
[receipt] jev: noul=0.94
[receipt] verdict: confirmsThe five verdicts and when each is emitted are owned by
rules/receipt-mapping.md.
null (ran, no support) and unobtainable (could not run) are distinct and
never collapsed — the same rule verification-receipt.md enforces.
Jev is a TypeSafe System One model: it takes structured natural-language
state and returns a typed answer with a probability (Noul / Choice /
Score), not generated prose.
That is exactly the shape a test assertion needs — a decision code can gate on,
with calibrated confidence.
Jev evaluates text, not images, so the browser side must feed it an
accessibility snapshot or page text; a screenshot is never a valid input.
See rules/state-extraction.md.
Run these three steps in order; each has a gate.
| Step | Name | Rule file | Gate |
|---|---|---|---|
| 0 | Provenance guard | rules/provenance.md | Expectation is a user-observable OUTCOME, not a restatement of the state text |
| 1 | State extraction | rules/state-extraction.md | Page TEXT state captured (a11y tree / page text); no screenshot |
| 2 | Jev call + mapping | rules/receipt-mapping.md | Jev primitive resolved and mapped to exactly one receipt verdict |
Callers supply two things — the page text state and the expectation — and receive one receipt.
<expectation> — a user-observable outcome in plain language, phrased as
something a person looking at the screen could confirm.
Provenance (Step 0) rejects an expectation that merely restates the captured
text (assertion-by-construction).--state-file <path> — a file holding the captured text state.
When omitted, the caller passes the state inline (the spec runners do this).--threshold-high N / --threshold-low N — override the default
decision bands (see rules/receipt-mapping.md); both are calibration knobs,
not magic numbers.The concrete call path is the committed, zero-dependency script — it builds the
Noul question, POSTs it, and maps the probability to a verdict per
receipt-mapping.md:
node ${CLAUDE_SKILL_DIR}/scripts/jev-call.mjs \
--expectation "the user sees an order-confirmation number" \
--state-file <captured-state.txt> --target <page-or-spec-id>State also reads from stdin when --state-file is omitted (the runners pipe
captured page text in). The script follows the TypeSafe HTTP API contract
documented by the official typesafe@typesafe-ai skill — install that skill for
the API guidance, prompting patterns, and SDK; this skill owns only the
semantic contract (question template, thresholds, receipt mapping) and does not
fork the wire format. The API is keyed by the TYPESAFE_API_KEY environment
variable. When TYPESAFE_API_KEY is unset or the API is unreachable, the check
could not run — the script emits unobtainable, never a confirms /
contradicts guess. Run node ${CLAUDE_SKILL_DIR}/scripts/jev-call.mjs --self-test to verify the mapping offline.
THEN assertionThe primary consumer is the shared UI spec grammar
(spec-run-contract.md).
A spec author writes a semantic assertion with the semantic: prefix, mirroring
the existing network: form:
- WHEN {role: "button", name: "Place order"} is clicked
THEN semantic: the user sees an order-confirmation numberBoth runners (aw-tester
and aw-tester-chrome)
capture the page text state, call this skill, and fold its receipt into the
spec verdict: confirms → the assertion passes; contradicts / null → it
fails; ambiguous / unobtainable → the spec is inconclusive, never a
silent pass.
ui-verify reuses that grammar verbatim, so a semantic THEN flows through to
PR/preview verification with no extra wiring.
verify-behavior verdict and
the raw Jev probability; it never invents its own grading number —
confidence(code) owns any score.observe-run's assertion-provenance rule.unobtainable — a verdict about the tooling, never a guessed pass or fail.evals/evals.json.evals/triggers.jsonl covers should-trigger and adjacent near-miss
queries.scripts/eval/l1.mjs; the accept/reject provenance
decision is an enumerable judgment guarded by a golden set.unobtainable as a failure, or null as a pass.typesafe@typesafe-ai — jev-call.mjs follows its documented contract, it
does not replace the guidance.name / description validate; node ${CLAUDE_SKILL_DIR}/scripts/validate-skill.mjs
(run from create-skill) reports PASS.verification-receipt.md exactly.semantic: assertion form is wired into spec-run-contract.md,
specs.md.template, and both runners.CLAUDE.md and README.md.39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.