CtrlK
BlogDocsLog inGet started
Tessl Logo

review-test-failures

Analyze dotnet/maui PR failures across maui-pr, maui-pr-devicetests, and maui-pr-uitests against the latest five completed runs on the PR's target branch. Report only whether failures are PR-related, with failure links grouped by pipeline. When no current results exist, skip evaluation and request /azp run.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.github/skills/review-test-failures/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-sequenced instruction body with a copy-paste output template, exact data contracts, explicit validation gates, and thorough error-recovery branches. Its weaknesses are mild redundancy across sections and an all-inline structure whose cross-bundle reference cannot be verified from within the skill directory.

Suggestions

Consolidate the repeated no-results/missing-evidence rules into the 'Check for results first' section and reference it from the styling section instead of restating them.

Move the detailed variant-matching and comparison rules (e.g., the Handler Does Not Leak / FlyoutHeaderScroll examples) into a reference file under references/ and keep only the core rule in SKILL.md.

Copy or link maui-ci-facts.md inside the skill bundle so the ../../docs/maui-ci-facts.md reference resolves reliably wherever the skill is installed.

DimensionReasoningScore

Conciseness

The body is dense and free of concept explanations Claude already knows, but at ~250 lines it repeats several guards: the no-results/`/azp run` rules appear in both 'Check for results first' and the styling section, missing-evidence caveats are restated ('Results missing from the bundle are not proof...' vs 'Never treat missing current results as passing'), and the five-run selection rules are split across two sections that could be tightened. Mostly efficient, but noticeably more than 'minor' trimmable material, so it sits below the 4 anchor.

3 / 5

Actionability

Everything is executable: exact context.json field names (pr.headRefOid, pr.baseRefName, history.pipelines), exact ref formats ('refs/heads/net11.0'), a copy-paste-ready literal comment template with exact badge URLs and HTML entities, exact commands ('/azp run PIPELINE_NAME', 'add_comment exactly once'), and a four-row attribution table with evidence requirements per class.

5 / 5

Workflow Clarity

The sections form an explicit pipeline — check results first (a hard validation gate on context.json), read evidence, compare the latest five target-branch runs, classify via the evidence table, then emit and publish the comment once. Error-recovery branches are explicit (unreadable context → stop; unverified pipeline → Insufficient data with link; already-running pipeline → 'Pending; wait for this run'), and verification checkpoints are stated ('Verify current builds belong to the captured PR head', 'require complete results for expected Helix work items').

5 / 5

Progressive Disclosure

Well-organized sections with clearly signaled, one-level-deep references: the gatherer script exists in the bundle ('the trusted scripts/Gather-TestFailureContext.ps1 collected it; do not rerun the gatherer') and the MAUI CI facts reference is named with its relevant sections. Not a 5 because the maui-ci-facts.md link points outside the skill directory (../../docs/) where it cannot be verified as part of the bundle, and the ~60-line literal template plus the detailed comparison/classification rules are all inlined with no references/ layer for the longer policy material.

4 / 5

Total

17

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A highly specific, distinctive description that names exact pipelines, the five-run comparison window, the output contract, and the no-results fallback. Its single weakness is the missing explicit 'Use when...' trigger clause, which caps completeness and slightly narrows the natural trigger terms.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to review, triage, or analyze CI/test failures on a dotnet/maui pull request (the /review tests command)'.

Include the natural synonyms 'test failures' and 'CI failures' alongside 'PR failures' so users' phrasing matches the description without exact wording.

Optionally state the consumer ('maintainers triaging PR CI results') to sharpen the when-side of the description.

DimensionReasoningScore

Specificity

The description names multiple concrete actions with full domain coverage: 'Analyze dotnet/maui PR failures across maui-pr, maui-pr-devicetests, and maui-pr-uitests against the latest five completed runs on the PR's target branch', 'Report only whether failures are PR-related, with failure links grouped by pipeline', and 'skip evaluation and request /azp run'. Every action is concrete and names the exact pipelines, comparison window, and fallback; not a 4 because coverage has no gaps for this domain.

5 / 5

Completeness

The 'what' is explicit (analyze three named pipelines, report PR-relatedness with grouped links, skip and request /azp run when no results), but there is no 'Use when...' clause or equivalent trigger guidance; 'When no current results exist' is a behavioral condition, not a usage trigger, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Strong natural keywords include 'PR failures', 'PR-related', the three pipeline names, and '/azp run', which maintainers of this repo would actually say. Not a 5 because the most natural user phrases 'test failures', 'CI failures', or 'triage' are absent — it only uses 'PR failures'.

4 / 5

Distinctiveness Conflict Risk

It occupies a clear niche — 'dotnet/maui PR failures' with the three exact pipeline names and '/azp run' — so it is unmistakable from any other skill and carries minimal conflict risk.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
dotnet/maui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.