CtrlK
BlogDocsLog inGet started
Tessl Logo

test-quality-reviewer

Evaluate the quality and efficacy of existing tests by reviewing test code against source code. Use when the user asks to review tests, validate test quality, audit test suites, check test efficacy, or assess whether tests are testing things properly. Prioritizes real interactions over mocking and simulation.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-crafted, opinionated skill body: it adds only non-obvious project-specific judgment (mock-vs-real priorities, verdict rubric, ecosystem heuristics) and gives Claude an exact output format to follow. Minor improvements are possible by deduplicating the anti-patterns table against earlier sections and adding edge-case handling to the workflow (e.g., missing source code).

DimensionReasoningScore

Conciseness

The body is dense and efficient — tables, prioritized principles, and per-framework heuristics with no explanations of concepts Claude already knows. The anti-patterns cheat sheet, however, restates items already covered in the philosophy and language guidance (mocking own interfaces, `toHaveBeenCalledWith`, `page.route`), which keeps it below the lean anchor 5 but well above noticeably-verbose anchor 3.

4 / 5

Actionability

For an instruction-only skill the guidance is fully actionable: a per-test evaluation criteria table, a four-value verdict taxonomy with definitions, a worked example evaluation table with realistic rows, and per-ecosystem good/bad/smell heuristics with diagnostic key questions ('If the backend broke, would this test catch it?'). Per the rubric's scoring note, absence of code is not penalized when guidance is this concrete.

5 / 5

Workflow Clarity

Steps 1-5 are clearly sequenced (read tests, read source, evaluate, identify gaps, produce report) with defined inputs and outputs at each step. It lacks explicit checkpoints or feedback loops (e.g., what to do when the source under test cannot be located, or how to prioritize findings in the report), so it falls short of the explicit-validation anchor 5; no validation cap applies since this is a read/evaluate/report skill, not a destructive or batch operation.

4 / 5

Progressive Disclosure

No bundle files exist and none are needed for a skill of this size; the body is well-organized with clear sections (philosophy, workflow, output format, language guidance, anti-patterns). At ~137 lines it exceeds the under-50-line simple-skill exception, and the language/framework guidance could plausibly be split into reference files, so it lands at anchor 4 rather than 5.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capability statement, explicit and well-phrased trigger guidance in third person, and a differentiating philosophy clause. The only weaknesses are a couple of missing natural trigger variants and slight overlap with general code-review territory.

DimensionReasoningScore

Specificity

Lists several concrete actions — 'Evaluate the quality and efficacy of existing tests', 'reviewing test code against source code', 'Prioritizes real interactions over mocking' — with only minor coverage gaps (e.g., producing a gap-analysis report is unstated). It goes beyond the 1-2 actions of anchor 3 but is not comprehensive enough for anchor 5.

4 / 5

Completeness

Explicitly answers both 'what' ('Evaluate the quality and efficacy of existing tests by reviewing test code against source code') and 'when' with concrete trigger phrases in a dedicated 'Use when...' clause. This matches the anchor-5 pattern of the rubric's good example almost exactly, exceeding anchor 4 where the 'when' would be less explicit.

5 / 5

Trigger Term Quality

The 'Use when' clause covers natural phrases users would actually say: 'review tests, validate test quality, audit test suites, check test efficacy, or assess whether tests are testing things properly'. A few common variants are missing (e.g., 'test coverage', 'are my tests good'), so it lands at anchor 4 rather than a fully comprehensive 5.

4 / 5

Distinctiveness Conflict Risk

The test-quality/efficacy framing plus the real-over-mocking stance carves a clear niche, but 'review tests' and 'audit test suites' carry minor overlap risk with generic code-review skills. It is more distinct than anchor 3 yet not the minimal-conflict clarity of anchor 5.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mattermost/mattermost-ai-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.