CtrlK
BlogDocsLog inGet started
Tessl Logo

review-testing

Review a code change for untested branches and error paths, tests that do not assert behavior, mirror tests that miss the source of truth, test-only production seams, duplicate coverage, brittle or nondeterministic tests, and behavior changes with no test changes. Use when reviewing tests, test coverage, test quality, or whether the tests prove the code works.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Excellent instruction-only skill body: a precise, dense checklist of defect patterns with per-item decision tests, a safe mutation-testing procedure, clear reporting thresholds, and no wasted tokens. The only gap is an explicit self-verification step for reported findings, which keeps workflow clarity at 4.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence throughout — it references Kent Beck's test desiderata by name, defines each defect class in one dense sentence, and never explains basic testing concepts. Every token earns its place, matching the anchor-5 example of lean, efficient content.

5 / 5

Actionability

As an instruction-only skill, the guidance is concrete and directly executable: each scope item carries its own decision test ('If the source script changes but the expected array does not, does the test fail?'), the Method gives an exact procedure ('Trace each new branch and error path in the change to at least one test that exercises it'), and Reporting specifies what each finding must state. No pseudocode or missing key details, so it sits at anchor 5 rather than 4.

5 / 5

Workflow Clarity

The section order (Scope → Method → Threshold → Reporting) forms a coherent review sequence, and the mutation-testing paragraph includes an explicit safety guard ('do so only in an isolated worktree... Never mutate a shared checkout') with a correctness condition (HEAD equals the reviewed commit). It falls short of anchor 5 because there is no explicit verification checkpoint on the reviewer's own findings, such as confirming each inferred gap by naming the test that should cover it — matching anchor 4's 'minor validation gaps'.

4 / 5

Progressive Disclosure

The skill is a single well-organized file under 50 lines with clear sections (Scope, Method, Threshold, Reporting) and no need for external references, which per the rubric's simple-skill guideline earns a 5 on well-organized sections alone. The AGENTS.md/CLAUDE.md mention is external project context, not a buried bundle reference.

5 / 5

Total

19

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it precisely enumerates what the skill detects, includes an explicit 'Use when...' trigger clause, and stays concrete without padding. The only room for improvement is broader natural trigger phrasing and slightly sharper differentiation from general code-review skills.

Suggestions

Add common user phrasings to the when-clause, e.g. 'flaky tests', 'unit tests', 'coverage gaps', or 'do my tests cover this change', to lift trigger-term coverage toward comprehensive.

Sharpen the opening so the lens is unambiguously about test quality rather than general change review (e.g. 'Review the tests in a code change for...' instead of 'Review a code change for...') to reduce overlap with general code-review skills.

DimensionReasoningScore

Specificity

The description enumerates a comprehensive catalog of concrete capabilities: 'untested branches and error paths, tests that do not assert behavior, mirror tests that miss the source of truth, test-only production seams, duplicate coverage, brittle or nondeterministic tests, and behavior changes with no test changes'. Every major defect class in the test-review domain is explicitly named, matching the comprehensive-coverage anchor rather than the minor-gaps anchor one level below.

5 / 5

Completeness

Both questions are answered explicitly and concretely: the first sentence is a detailed 'what', and 'Use when reviewing tests, test coverage, test quality, or whether the tests prove the code works' is an explicit 'when' with concrete trigger phrases. This matches the anchor-5 example pattern exactly.

5 / 5

Trigger Term Quality

The when-clause 'Use when reviewing tests, test coverage, test quality, or whether the tests prove the code works' provides good natural keyword coverage. It falls short of anchor 5 because several common user phrasings are absent ('flaky tests', 'unit tests', 'coverage gaps'), but is well above anchor 3 since multiple natural variations are present.

4 / 5

Distinctiveness Conflict Risk

The testing-review niche is clear and triggers are distinct ('reviewing tests, test coverage, test quality'), but the opening 'Review a code change for...' creates minor overlap risk with general code-review skills that could also claim test-related review requests. Mostly distinct with minor overlap against a closely related skill family, matching anchor 4 rather than the minimal-conflict anchor 5.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
perihelionhq/perihelion-platform-context
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.