CtrlK
BlogDocsLog inGet started
Tessl Logo

pr-test-checker

Grade whether a Positron PR has adequate test coverage. Evaluates new tests in the PR, checks existing coverage for changed source, and suggests concrete additions when coverage is insufficient. Used by the pr-test-checker GitHub Action.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/pr-test-checker/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-sequenced grading workflow whose repo-specific knowledge (taxonomy, surfaces, tags) genuinely earns its tokens. Its weaknesses are repetition of the e2e cost guidance across sections and a monolithic ~200-line body that neither moves hotspot/template detail into reference files nor mentions the bundle's own context-gathering scripts.

Suggestions

Move the Windows/web hotspot lists and the full output-format template into a reference file (e.g., references/report-template.md), keeping only the verdict table and section names inline.

Deduplicate the e2e cost guidance — state it once in Cost guidance and have Constraints reference it in one line instead of restating the rule.

Mention the bundle's scripts (scripts/gather-pr-context.mjs, scripts/gather-local-context.mjs) and how their context.json relates to the "Inputs you'll receive" section so the bundle structure is discoverable.

DimensionReasoningScore

Conciseness

Mostly efficient — the test taxonomy, surface tables, and tag mechanics are genuine repo-specific knowledge Claude doesn't have — but the e2e cost guidance is repeated almost verbatim in both "Cost guidance" and "Constraints", and the output template embeds long editorial parentheticals that could be trimmed. It is not a 4 because the duplication and padded sections are noticeable, and not a 2 because most content is non-obvious domain knowledge.

3 / 5

Actionability

Fully executable throughout: copy-paste grep commands (e.g., `grep -r "from.*<filename>" src/vs/**/test/`), exact path patterns per runner, a precise verdict table with fixed emojis, a complete output markdown template, and a two-case decision rule. Every instruction names the concrete file, command, or output line to produce.

5 / 5

Workflow Clarity

The seven investigation steps are explicitly ordered with validation checkpoints (step 4 requires reading candidate tests to confirm they exercise the changed behavior; step 5 conditions on the test surface; step 7 requires grep-confirming a scenario exists before suggesting it), plus a tool-call budget with a defined fallback behavior ("lean toward Insufficient with a note") — matching the anchor for clear sequence with explicit validation and error-recovery handling.

5 / 5

Progressive Disclosure

Sections are well organized and external references (CLAUDE.md, .claude/rules/vitest-tests.md, test/e2e/infra/test-runner/test-tags.ts) are clearly signaled at one level deep, but the body is a ~200-line monolith: the Windows/web hotspot lists and the full output template are inline content that plausibly belongs in reference files, and the three bundle scripts in scripts/ are never mentioned or navigated to from the body. It is not a 4 because the bundle structure itself is undiscoverable and substantial separable content is inlined.

3 / 5

Total

16

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-scoped, third-person description with concrete actions and a distinct niche. Its main weakness is the absence of explicit "when to use" trigger phrasing, which both caps completeness and leaves trigger-term coverage thinner than the domain's natural vocabulary.

Suggestions

Add an explicit trigger clause, e.g., "Use when reviewing a Positron pull request's test coverage or when a PR lacks tests for changed source".

Include the natural synonyms users would say — "unit tests", "regression test", "e2e", "pull request" — so the skill triggers on real phrasings.

DimensionReasoningScore

Specificity

The description lists four concrete, distinct actions — "Grade whether a Positron PR has adequate test coverage", "Evaluates new tests in the PR", "checks existing coverage for changed source", "suggests concrete additions when coverage is insufficient" — which comprehensively covers the skill's behavior, matching the anchor for multiple specific concrete actions.

5 / 5

Completeness

The "what" is clear and concrete, but there is no "Use when..." clause or equivalent trigger guidance; "Used by the pr-test-checker GitHub Action" states deployment context, not when Claude should invoke it, so the missing-trigger cap applies. It is not a 2 because the "what" is fully explicit.

3 / 5

Trigger Term Quality

Natural terms like "test coverage", "PR", "tests", and "Positron" are present and would be said by a user, but common synonyms users might invoke ("unit tests", "regression test", "e2e", "pull request") are missing, placing it just below the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

"Grade whether a Positron PR has adequate test coverage" carves out a clear niche (Positron PRs, test-coverage grading) with distinct triggers and minimal overlap with generic testing or review skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.