CtrlK
BlogDocsLog inGet started
Tessl Logo

review-tests

Use when checking whether a branch is adequately tested — building and running the affected workspaces, then auditing changed code for missing test files, untested exports, untested error paths and untested routes — and reporting gaps with T-M/U/E/R/S IDs in the four-field format.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced procedure with strong validation loops — the report-skeleton-first design and shape checks are exemplary. Its weaknesses are repetition of the shape rules (stated three times) and a monolithic single-file layout where the detailed audit sub-procedures could live in reference files.

Suggestions

State the five-heading shape rule and the forbidden-heading list once (Step 0) and have later checks reference it, removing the duplicated grep command and the repeated "## Summary / ## Verdict / ## Scope and method / ## Bottom line" enumeration.

Move the credential-enumeration sub-procedure and the tautological-test shape checklist into reference files (e.g. references/audit-patterns.md) linked one level deep from Step 4, shortening SKILL.md to the workflow itself.

Trim rhetorical framing (e.g. "however good the audit inside it", "it is the argument the finding rests on") to bare imperatives — the rules land without the emphasis.

DimensionReasoningScore

Conciseness

The body is mostly dense and earned, but the five-heading shape rule and the forbidden-heading list ("`## Summary`, `## Verdict`, `## Scope and method` and `## Bottom line`") are repeated three times, the `grep -c` shape check appears twice, and rhetorical asides ("however good the audit inside it", "and that is the one way this run fails outright") pad the prose — more than the minor trimming of anchor 4, so it sits at anchor 3.

3 / 5

Actionability

Guidance is fully executable throughout: exact commands (`git diff --name-only "$BASE"...HEAD`, `bun run --cwd <workspace-path> test:run`, the verbatim `TEST-REVIEW.md` heredoc, the two `grep -c` checks), an exact four-field finding template, and concrete test patterns (`Effect.flip`, `createXxxRoutes(createTestLayer())`) covering the common cases — matching anchor 5 rather than the minor-gaps anchor 4.

5 / 5

Workflow Clarity

Steps 0–4 are clearly sequenced with explicit validation checkpoints (the shape check that "must print `5`", run "as soon as the first gap is in the file", and the two-count final check) and error-recovery loops ("No step is a stop"; restore the skeleton if a count is low) — the anchor 5 pattern of sequence plus validation plus feedback, not merely the most-checkpoints-present anchor 4.

5 / 5

Progressive Disclosure

Sections are well-organized with clear `##` headers and the two `wiki/conventions/` pointers are clearly signaled inline, but there is no bundle structure at all — long sub-procedures such as the credential-enumeration walkthrough and the tautological-test checklist are inlined in a ~215-line file where anchor 5 would split them into one-level-deep reference files; better than anchor 3, whose structure and signaling would be genuinely unclear.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states an explicit "Use when" trigger, enumerates concrete capabilities comprehensively, and is unmistakably distinct via its T-M/U/E/R/S reporting vocabulary. The only weakness is trigger-term breadth — natural synonyms like "test coverage" or "missing tests" are not present.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "building and running the affected workspaces", "auditing changed code for missing test files, untested exports, untested error paths and untested routes", "reporting gaps with T-M/U/E/R/S IDs in the four-field format" — giving comprehensive coverage of what the skill does, matching the anchor 5 example rather than the minor-gaps anchor 4.

5 / 5

Completeness

It explicitly answers both questions: the "what" is building/running tests and auditing changed code for four named gap classes, and the "when" is stated as an explicit "Use when checking whether a branch is adequately tested" trigger — clearly above anchor 4, where the "when" would be less explicit.

5 / 5

Trigger Term Quality

"checking whether a branch is adequately tested" is a phrase a user would naturally say, and "branch", "tested", "workspaces" are relevant terms, but common synonyms such as "test coverage", "missing tests", or "unit tests" are absent — good keyword coverage with a few natural terms missing, above anchor 3 but short of the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

The description carves a clear niche — a branch-level test audit producing a report with a named T-M/U/E/R/S taxonomy in a four-field format — which is unmistakable against generic run-the-tests or code-review skills; not anchor 4, since even minor overlap risk is suppressed by the distinctive output contract.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
englishstreetventures/osn
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.