CtrlK
BlogDocsLog inGet started
Tessl Logo

qa-review

QA review for code changes — test coverage analysis, edge case identification, test plan generation, regression detection, test health tracking over time.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

80%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, actionable QA-review skill with concrete templates and a clear sequenced methodology, held back by the absence of explicit validation checkpoints in its workflow and no use of file-based progressive disclosure despite a moderately long single-file body.

Suggestions

Add an explicit validation/verification checkpoint to the review methodology — e.g. 'Confirm each coverage gap against the actual diff before reporting it' — to introduce a feedback loop and lift workflow clarity to 3.

Consider moving the detailed output-format template or edge-case taxonomy into a one-level-deep referenced file (e.g. OUTPUT_FORMAT.md) linked from the overview, to demonstrate progressive disclosure and keep SKILL.md a concise entry point.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence — it never explains concepts Claude already knows (what a test is, how coverage works) and every section (methodology, edge-case taxonomy, output template, fix-first model) earns its place, matching the score-3 'every token earns its place' anchor.

3 / 3

Actionability

It provides copy-paste-ready output and test-plan templates with concrete placeholders, an explicit edge-case checklist, and specific signal field names (obligation_type: testing, immediacy: batch), giving concrete executable guidance rather than vague direction; absence of code is fine for this instruction-only skill.

3 / 3

Workflow Clarity

The review methodology is a clearly numbered 5-step sequence, but there are no explicit validation/verification checkpoints or feedback loops in the process (e.g. 'confirm each gap against the actual diff before reporting'), so it sits at the score-2 anchor of sequence present but checkpoints missing or implicit rather than the score-3 anchor requiring explicit validation steps.

2 / 3

Progressive Disclosure

The skill is a single ~100-line self-contained file with no external references; it is well-sectioned and easy to navigate, but it does not demonstrate the overview-pointing-to-details pattern and exceeds the 50-line simple-skill threshold where single-file organization alone earns a 3, so it does not reach the score-3 anchor.

2 / 3

Total

10

/

12

Passed

Description

82%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, third-person description with strong natural trigger terms and a clear testing niche, weakened only by the absence of an explicit 'Use when...' clause that would fully answer when to invoke it.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when reviewing code changes before merge, or when the user asks about test coverage, edge cases, or regression risks', to satisfy the 'when should Claude use it' requirement and lift completeness to 3.

Consider adding common natural phrasings users actually say, such as 'what tests am I missing' or 'is this safe to merge', to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

The description enumerates five concrete actions — 'test coverage analysis, edge case identification, test plan generation, regression detection, test health tracking over time' — matching the score-3 anchor of listing multiple specific concrete actions.

3 / 3

Completeness

It clearly states what the skill does but lacks any 'Use when...' clause or equivalent explicit trigger guidance, so per the judging guidelines completeness is capped at 2; it is not a 3 because the 'when' is only implied, not explicit.

2 / 3

Trigger Term Quality

It surfaces natural terms a user would actually say — 'QA review', 'test coverage', 'test plan', 'edge case', 'regression', 'test health' — giving good coverage rather than technical jargon.

3 / 3

Distinctiveness Conflict Risk

The QA/testing niche with distinct triggers (test coverage, edge cases, regression, test health) is clearly separated from general code-review concerns and unlikely to fire for the wrong skill.

3 / 3

Total

11

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
nearai/ironclaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.