CtrlK
BlogDocsLog inGet started
Tessl Logo

verify-tests-fail-without-fix

Verifies tests catch the bug. Auto-detects test type (UI tests, device tests, unit tests) and dispatches to the appropriate runner. Supports two modes - verify failure only (test creation) or full verification (test + fix validation).

59

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.github/skills/verify-tests-fail-without-fix/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable and operationally safe: executable commands throughout, a single-script ownership model with explicit blocked-state handling, and a report contract that removes all ambiguity from the inverted pass/fail semantics. Its weaknesses are duplication — the modes, inverted semantics, and output layout are each stated three to four times — and a monolithic 265-line structure where troubleshooting and parameter reference material would be better split into a one-level-deep reference file. Both issues are about token efficiency, not correctness.

Suggestions

Collapse the mode duplication: keep the "Mode 1"/"Mode 2" sections with their commands and delete the redundant restatements in "Requirements" and "What It Does", which repeat the same fail-without-fix / pass-with-fix logic almost verbatim.

State the inverted pass/fail semantics once (the table plus one warning sentence); remove the repeated reminders in Step 3, Step 4, and the intro, trusting the initial statement to carry.

Move the Output Files directory tree, Troubleshooting table, and Optional Parameters into a one-level-deep reference file (e.g. references/parameters.md) and link to it, shrinking SKILL.md to the workflow core.

DimensionReasoningScore

Conciseness

The body is mostly project-specific knowledge Claude could not know (runner scripts, paths, parameters, output contracts), which earns its tokens — but the same facts are repeated several times: the two modes are explained in "Workflow Step 1", again in "Mode 1"/"Mode 2", again in "Requirements", and again in "What It Does"; the inverted pass/fail semantics appear four times (the table, "NEVER say 'verification passed'", the Step 3 reminder, and the Step 4 bullet list); and the output-file layout is described twice ("Output Files" table and the example tree). This matches anchor 3 — mostly efficient with unnecessary duplication that could be tightened — rather than 2, since almost no space is spent on concepts Claude already knows.

3 / 5

Actionability

Every instruction is copy-paste executable: exact pwsh invocations with real parameter values ("pwsh .github/skills/verify-tests-fail-without-fix/scripts/verify-tests-fail.ps1 -Platform android -TestFilter 'Maui12345' -RequireFullVerification"), the referenced script actually exists in the bundle (scripts/verify-tests-fail.ps1), the expected output markers are shown verbatim, and the Optional Parameters section gives concrete syntax for every flag including the frozen-fixture "$(git rev-parse HEAD^)" variant. Not below 5: the common cases (auto-detect, explicit filter, full verification) are all covered with working commands.

5 / 5

Workflow Clarity

The four-step workflow is clearly sequenced with strong validation checkpoints: mode determination from the git diff, a single script owning all fix-file reversion (with an explicit prohibition on manual git cleanup), unambiguous terminal markers (VERIFICATION PASSED / FAILED / error-timeout → Blocked), explicit background-session handling, and a report contract that enumerates what each marker means in each mode. The operation is destructive (reverting fix files), but the guardrails — the script owning transitions, reporting Blocked rather than cleaning up, and the troubleshooting table mapping failure → cause → solution — provide exactly the feedback loops the top anchor requires.

5 / 5

Progressive Disclosure

The body is well-sectioned with clear headers (Supported Test Types, Workflow, Expected Output, Output Files, Troubleshooting, Optional Parameters) and its one script reference is real and matches the bundle layout (scripts/verify-tests-fail.ps1). It does not reach 5 because the file is a ~265-line monolith: content that would sit naturally in a one-level-deep reference — the output-file directory tree, the troubleshooting table, and the full parameter reference — is inlined in SKILL.md, and no reference files exist in the bundle. It is comfortably above 3 since what is present is organized and navigable.

4 / 5

Total

17

/

20

Passed

Description

55%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid, concrete description with good third-person specificity and a clear two-mode decomposition, but it completely lacks a "when to use" trigger clause, which caps both completeness and trigger-term quality. It also under-communicates its most distinctive behavior (tests must fail without the fix) in favor of generic test-type vocabulary. Adding an explicit "Use when..." sentence covering test-first development and fix validation would lift two dimensions at once.

Suggestions

Append an explicit trigger clause, e.g. "Use when writing regression tests for a bug, practicing test-first/TDD, or validating that a fix makes the failing tests pass" — this directly addresses the completeness cap and adds natural trigger terms.

Include user-facing synonyms for the core mechanism, such as "reproduce the bug", "regression test", and "tests fail without the fix", which are the phrases a user would naturally say when they need this skill.

State the differentiator explicitly ("confirms tests FAIL before the fix is applied") so the description cannot be confused with a generic run-the-tests skill.

DimensionReasoningScore

Specificity

The description names several concrete actions — "Auto-detects test type (UI tests, device tests, unit tests) and dispatches to the appropriate runner" and "Supports two modes - verify failure only (test creation) or full verification (test + fix validation)" — which are specific and grounded in the skill's actual behavior. It falls short of a 5 because the headline "Verifies tests catch the bug" leaves the core mechanism (running tests without the fix to confirm they fail) implicit, and it is above a 3 because it lists multiple concrete actions rather than 1-2.

4 / 5

Completeness

The "what" is clear and well-decomposed (auto-detection, dispatch, two modes), but there is no "Use when..." clause or equivalent explicit trigger guidance anywhere in the description. Per the judging guidelines, a missing 'Use when...' clause caps completeness at 3, which this description hits exactly: clear what, no when.

3 / 5

Trigger Term Quality

Relevant keywords are present ("tests", "bug", "UI tests", "device tests", "unit tests", "verification", "fix validation"), but common natural variations users would say are missing: "reproduce the bug", "regression test", "test-first", "TDD", "red phase", or "tests fail without the fix". It sits between anchor 3 (some relevant keywords, missing common variations) and anchor 4 — the domain vocabulary is right but the user-facing synonyms are absent, so 3 is the best fit.

3 / 5

Distinctiveness Conflict Risk

The core purpose (verifying tests catch a bug before/with a fix) is a distinct niche, but the description leans on generic test vocabulary ("UI tests, device tests, unit tests", "runner") that overlaps heavily with plain test-execution skills, and it never states the differentiator — verifying failure without the fix — in the trigger surface. It is more specific than anchor 2's "very broad, high overlap" but does not reach anchor 4's "mostly distinct", so 3 fits best.

3 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dotnet/maui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.