CtrlK
BlogDocsLog inGet started
Tessl Logo

test-evidence-review

Quality review of test files and evidence — goes beyond existence, evaluates assertion coverage. ADEQUATE/INCOMPLETE/MISSING/NOT ASSESSED per story.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/test-evidence-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually well-engineered procedural skill: the workflow is fully sequenced with real validation guards, the guidance is copy-paste concrete, and edge conditions (empty scope, unreadable evidence, undeterminable story type) each have an explicit home in the verdict vocabulary. The main cost is token weight — several multi-line design-rationale essays and the sheer length of inline per-type rules could be trimmed or moved to a reference file.

Suggestions

Cut the meta-commentary blockquotes ("Why this needed saying", "This is a recurring shape: ... what does this emit when the set is empty?") — they argue for the skill's design to a reader, not to the executor running it.

Move the per-story-type gate resolution rules (qa.level / testing.strict matrix) and the full report template into a bundled reference file, keeping SKILL.md as the sequenced overview.

Compress the verdict-ranking paragraph in Section 6 into a one-line ordering rule (e.g. "MISSING > INCOMPLETE > NOT ASSESSED > ADEQUATE > WAIVED") plus a single sentence on why, rather than the current multi-paragraph justification.

DimensionReasoningScore

Conciseness

The body is mostly efficient procedural instruction, but it carries several clearly unnecessary explanation blocks: the Section 2 blockquote essay ("This is a recurring shape: the sophisticated inner rule present, the outer boundary unguarded... what does this emit when the set is empty?") and Section 6's "Why this needed saying" rationale passages are skill-design commentary, not execution guidance. This matches anchor 3 (mostly efficient, some unnecessary explanation that could be tightened) — above anchor 2 because the padding is confined to a few blockquotes rather than pervasive, below anchor 4 because those passages run to tens of lines.

3 / 5

Actionability

The guidance is copy-paste ready throughout: exact Grep invocations with pattern, glob, output_mode and -A flags (Section 2); concrete per-engine test roots (`tests/unit/[system]/`, `Assets/Tests/EditMode/`, `Source/<Module>/Private/Tests/`); numeric thresholds ("3+ assertions per test function → normal"); per-engine naming patterns; and a complete fill-in report template. This matches anchor 5 (fully executable, specific examples covering the common cases) — nothing is pseudocode or hand-waved.

5 / 5

Workflow Clarity

The seven sections form a clear, sequenced workflow with explicit checkpoints and error-recovery routing: the zero-stories guard stops execution and routes to the correct next skill; the stated-evidence-path-first rule with a documented fallback search; full-read fallback when a Test Evidence section is missing or ambiguous; and the ask-before-write confirmation for the optional report. This matches anchor 5 (clear sequence, explicit validation steps, feedback loops); the read-only nature of the skill means no destructive-operation cap applies.

5 / 5

Progressive Disclosure

No bundle files exist, and the body's references (`.claude/docs/coding-standards.md`, `.claude/rules/test-standards.md`, `.claude/docs/directory-structure.md`, `.claude/docs/automation-modes.md`) are one level deep and clearly signaled at the point of use, matching anchor 4 (good structure, references mostly clear, minor organization gaps). It is not anchor 5 because at ~350 lines the SKILL.md is a monolith — the detailed per-story-type gate logic and the full report template are candidates for a bundled reference file — though the references that do exist are external project docs rather than nested skill files, so it is comfortably above anchor 3.

4 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a distinct, specific capability but reads as half-written: the 'when' half of the what/when pair is entirely absent and the trigger vocabulary is narrow. Adding an explicit 'Use when...' clause with natural QA/review phrases would move completeness and trigger quality up a level.

Suggestions

Add an explicit trigger clause, e.g. "Use when reviewing test quality before QA hand-off, story closure, or milestone audit, or when a story's test or evidence quality is in question."

Broaden trigger terms to natural user phrasings: "test quality", "QA", "sign-off", "evidence", "coverage gaps" — not just "test files" and "assertion coverage".

Fold one or two more concrete capabilities into the 'what' (sign-off completeness, screenshot/artefact checks) so the description covers the skill's actual scope rather than only assertion coverage.

DimensionReasoningScore

Specificity

The description names its domain ("test files and evidence") and 1-2 concrete actions ("Quality review", "evaluates assertion coverage"), but does not enumerate the fuller set of capabilities the body actually performs (sign-off checks, screenshot/artefact completeness, criterion linkage). This matches anchor 3 (domain + 1-2 concrete actions, not comprehensive); it is above anchor 2 because the actions are specific rather than generic, and below anchor 4 because the verdict vocabulary line describes output labels rather than additional concrete actions.

3 / 5

Completeness

The 'what' is clear ("Quality review of test files and evidence — goes beyond existence, evaluates assertion coverage"), but there is no 'Use when...' clause or any equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Anchor 3 fits: clear 'what', 'when' entirely missing — not anchor 4, which requires an explicit (if imperfect) 'when'.

3 / 5

Trigger Term Quality

Relevant keywords are present ("test files", "evidence", "assertion coverage", "test"), but common natural variations users would say — "test quality", "QA", "sign-off", "coverage" in the generic sense — are absent, matching anchor 3 (some relevant keywords, missing variations/synonyms). It is above anchor 2 because the terms present are domain-natural rather than generic, and below anchor 4 because the natural-phrase coverage is thin for a QA-adjacent skill.

3 / 5

Distinctiveness Conflict Risk

The pairing of "evidence" with "assertion coverage" and the ADEQUATE/INCOMPLETE/MISSING/NOT ASSESSED verdict vocabulary carve a clear niche that distinguishes it from generic code-review or test-running skills, matching anchor 4 (mostly distinct, minor overlap risk with closely related skills such as a smoke-check or story-closure skill). It is not anchor 5 because the first half ("Quality review of test files") alone would overlap with ordinary test/CI review requests.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
Donchitos/Claude-Code-Game-Studios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.