CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-pr-tests

Evaluates tests added in a PR for coverage, quality, edge cases, and test type appropriateness. Checks if tests cover the fix, finds gaps, and recommends lighter test types when possible. Prefer unit tests over device tests over UI tests. Triggers on: 'evaluate tests in PR', 'review test quality', 'are these tests good enough', 'check test coverage', 'is this test adequate', 'assess test coverage for PR'.

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, highly actionable evaluation workflow with concrete commands, tables, a decision tree, and a complete output template, all grounded in project-specific conventions rather than general knowledge. Its main weaknesses are moderate: the Output Format template and criteria detail could be trimmed or split into reference files, and the workflow lacks an explicit validation checkpoint on the script's output.

Suggestions

Move the nine detailed evaluation criteria (or at least the long convention/flakiness tables) into a references/ file (e.g., references/criteria.md) and keep one-line summaries in SKILL.md, shortening the always-loaded body.

Condense the ~70-line Output Format template to a short skeleton showing section order and verdict labels, letting the full template live in a reference file.

Add an explicit checkpoint after Step 1, e.g., 'Verify CustomAgentLogsTmp/TestEvaluation/context.md exists and lists fix files; if not, apply the Troubleshooting table before continuing', to close the workflow validation gap.

DimensionReasoningScore

Conciseness

The ~330-line body is dense with project-specific material Claude cannot infer (MAUI test conventions, project names, anti-pattern tables, preference order) and avoids explaining general concepts, fitting the 'efficient; minor instances that could be trimmed' anchor. It is not a 5 because the 70-line Output Format template and some 'How to check' bullets (e.g., 'Read the fix files to understand: What changed / Why it changed') could be tightened, and not a 3 because there is no genuinely unnecessary explanation.

4 / 5

Actionability

Copy-paste-ready commands ('pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1 -BaseBranch "origin/main"'), a concrete decision tree, explicit project-file mappings ('*.UnitTests.csproj', 'TestCases.Shared.Tests'), a full report template, and good/bad code examples make the guidance fully executable. Score 4 would require missing key details for common cases, which is not evident.

5 / 5

Workflow Clarity

Steps 1-4 are clearly sequenced (gather context via script, understand the fix, evaluate against all criteria, produce the report) and the Troubleshooting table provides error recovery, fitting 'clear sequence with most checkpoints present'. It is not a 5 because there is no explicit checkpoint verifying the script's report was produced before proceeding (recovery is reactive via the troubleshooting table), though the destructive-operation cap does not apply since this skill is read-only analysis.

4 / 5

Progressive Disclosure

Structure is good: well-organized sections, a single one-level-deep bundle reference (scripts/Gather-TestContext.ps1, verified to exist) invoked consistently, and automated checks correctly delegated to the script rather than inlined. It is not a 5 because the nine detailed evaluation criteria (~200 lines of tables and examples) live entirely inline in SKILL.md where a references/ split would shorten the always-loaded body, though the inline form remains navigable.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: third-person voice, multiple concrete capabilities, an explicit 'Triggers on' clause with natural user phrasings and synonyms, and a distinct niche. It closely matches the good_overall_examples pattern and earns top marks on all four dimensions.

DimensionReasoningScore

Specificity

"Evaluates tests added in a PR for coverage, quality, edge cases, and test type appropriateness. Checks if tests cover the fix, finds gaps, and recommends lighter test types" lists multiple specific concrete actions covering the skill's full scope, matching the anchor for comprehensive coverage. It is above the level-4 anchor because there are no meaningful gaps in the action list, and it is not below since nothing is generic.

5 / 5

Completeness

The description explicitly answers both 'what' (evaluates coverage, quality, edge cases, test type appropriateness; finds gaps; recommends lighter test types) and 'when' ("Triggers on: 'evaluate tests in PR', ...") with concrete trigger phrases. This is exactly the level-5 anchor; level 4 would apply only if the 'when' were less explicit.

5 / 5

Trigger Term Quality

The trigger list ('evaluate tests in PR', 'review test quality', 'are these tests good enough', 'check test coverage', 'is this test adequate', 'assess test coverage for PR') covers natural user phrasings plus synonyms (evaluate/review/check/assess) of the same intent. This matches the comprehensive-synonyms anchor; score 4 would require notable missing natural terms, which is not the case here.

5 / 5

Distinctiveness Conflict Risk

"Evaluates tests added in a PR" with test-focused triggers carves out a clear niche distinct from general code-review or file-processing skills, with minimal conflict risk. It is not a 4 because the trigger phrases ('are these tests good enough', 'is this test adequate') are specific to this skill and would not plausibly fire a different one.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dotnet/maui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.