CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-review

Review whole-repo test quality, rerun coverage, score remaining worth-testing files, inspect slow-drift and stale test debt, and publish the next testing batch. Use every few weeks or before large breaking changes and rearchitecture.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a dense, well-sequenced audit workflow with runnable coverage and debt-scanning commands and a clear audit-only boundary. Weaknesses are mild: some duplication of stop-condition guidance, an ambiguous coverage-directory placeholder, no concrete scoring method for ranking files, unspecified artifact filenames, and implicit rather than explicit verification checkpoints.

Suggestions

Consolidate the stop-condition guidance (Goal, Core Rules, and Stop Conditions all restate it) into a single section to trim redundant tokens.

Give step 3 a concrete scoring method or rubric (e.g., a point scale per seam type) so file ranking is reproducible, and specify exact output file paths and names for step 4's artifacts.

Clarify the '.coverage-repo-YYYY-MM-DDx' placeholder (what the trailing 'x' means) and add an explicit checkpoint that lcov.info was produced before proceeding to scoring.

DimensionReasoningScore

Conciseness

The body is lean and telegraphic with zero concept re-teaching, but stop-condition guidance is stated three times ('stop fake work before it starts' in Goal, 'If the remaining misses are mostly low-ROI dust, say stop' in Core Rules, and the Stop Conditions section). It is not 5 because that duplicated stop guidance could be trimmed; not 3 because there is no real over-explanation or padding.

4 / 5

Actionability

Steps 1 and 2 give copy-paste executable commands ('bun test --coverage --coverage-reporter=lcov ...', 'pnpm test:profile -- --top 25', concrete rg patterns), but gaps remain: the '.coverage-repo-YYYY-MM-DDx' directory template is ambiguous, step 3 gives scoring criteria with no scoring method or formula, and step 4's artifact filenames ('a markdown map', 'a package TSV') are unspecified. It is not 5 because of those missing executable specifics; not 3 because the commands that are present are genuinely runnable.

4 / 5

Workflow Clarity

Five numbered steps are clearly sequenced with a capture list in step 1 and per-step expectations, and the audit-first scope is explicit. Verification is mostly implicit — there is no checkpoint confirming lcov.info was actually produced before scoring, and no verification of written artifacts — so it is not 5; it is not 3 because the sequence, inputs, and outputs per step are well defined.

4 / 5

Progressive Disclosure

There are no bundle files (no references/, scripts/, or assets/), and this ~170-line single-file skill is cleanly sectioned with its external inputs clearly signaled via '@.agents/rules/...' paths. Content is appropriately self-contained; minor gap in that the scoring rules and roadmap template could be split into a reference file as the skill grows. It is not 5 because it exceeds the small-skill case where sectioning alone would suffice; not 3 because nothing is buried or inlined that clearly belongs elsewhere.

4 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it enumerates concrete actions that mirror the workflow, uses third-person imperative voice, and pairs them with an explicit and specific 'Use when' trigger clause. Trigger-term synonym coverage and a slight overlap risk with general test/code-review skills are the only weaknesses.

Suggestions

Add natural trigger synonyms such as 'test audit' or 'test suite health' that a user would plausibly say when requesting this periodic review.

Sharpen distinctiveness by emphasizing the roadmap/batch-locking aspect (e.g., 'lock and maintain a testing roadmap') to separate it from one-off test-writing skills.

DimensionReasoningScore

Specificity

The description lists five concrete actions — 'rerun coverage', 'score remaining worth-testing files', 'inspect slow-drift and stale test debt', 'publish the next testing batch', 'Review whole-repo test quality' — each mapping to a distinct workflow step, matching the comprehensive multi-action anchor. It is not 4 because there are no meaningful gaps in the action coverage.

5 / 5

Completeness

It explicitly answers 'what' with five concrete actions and 'when' with the clause 'Use every few weeks or before large breaking changes and rearchitecture' — concrete trigger phrases for both. It is not 4 because the 'when' is already explicit and specific rather than merely present.

5 / 5

Trigger Term Quality

Good keyword coverage including natural phrases like 'test quality', 'coverage', 'test debt', 'breaking changes', and 'rearchitecture', but a few common user phrasings are missing (e.g., 'test audit', 'test suite', 'slow or flaky tests'). It is not 5 because synonym coverage is incomplete, and not 3 because the present terms are natural things a user would actually say.

4 / 5

Distinctiveness Conflict Risk

Terms like 'whole-repo test quality', 'worth-testing files', 'test debt', and 'testing batch' carve out a clear periodic testing-audit niche. It is not 5 because there is minor overlap risk with a generic test-writing or code-review skill triggered by 'review' and 'coverage'; not 3 because the batch/roadmap framing is clearly distinct from those.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
udecode/plate
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.