CtrlK
BlogDocsLog inGet started
Tessl Logo

run-device-tests

Build and run .NET MAUI device tests locally with category filtering. Supports iOS, MacCatalyst, Android on macOS; Android, Windows on Windows. Use TestFilter to run specific test categories.

62

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.github/skills/run-device-tests/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured body whose commands are immediately executable and whose heavy implementation scripts are properly externalized. Its main cost is redundancy — duplicated example commands and prerequisites sections — and inlined deep-dive material that inflates token usage without adding proportional value.

Suggestions

Delete the "Examples" section (or fold its two unique cases into "Scripts") — it repeats eight commands already shown verbatim above.

Merge "Prerequisites" into "Tools Required" — they list the same requirements (xharness, .NET SDK, Xcode, Android SDK, Windows SDK) twice.

Move the "How Test Filtering Works" platform table, "Copilot Gate" details, and example xharness invocations into a references/ file (e.g., references/test-filtering.md) and link to it from the body.

DimensionReasoningScore

Conciseness

The body contains substantial duplication: the "Examples" section repeats eight commands nearly verbatim from the "Scripts" section, and "Prerequisites" restates "Tools Required". There is also deep implementation detail (xharness flag-by-flag explanations, gate retry semantics) that could be tightened or moved out. It is not severely padded — most content carries real information — so it sits at anchor 3 rather than 2.

3 / 5

Actionability

Every command is copy-paste ready: concrete pwsh invocations with all parameters, exact project paths, artifact locations, executable xharness examples, and step-numbered troubleshooting with exact commands. The common cases (per-platform runs, filtering, build-only, rebuild) are all covered.

5 / 5

Workflow Clarity

The Workflow section gives a clear two-step sequence (run via script, check console/artifacts/log), device detection pipelines are documented per-platform, and Troubleshooting provides explicit feedback loops (rebuild → check "TestFilter:" in output → verify category case-sensitively). It falls short of anchor 5 because the main workflow does not itself state how to confirm success/failure (e.g., what a passing vs. failing run looks like) and relies on the troubleshooting section for recovery.

4 / 5

Progressive Disclosure

The 76KB Run-DeviceTests.ps1 and its test file are correctly externalized to scripts/ rather than inlined, and the body consistently points to them plus repo-side shared scripts. Structure is clear with well-organized headers. It does not reach anchor 5 because everything else is inlined in one ~310-line body — the Test Filtering internals, xharness invocation details, and category reference could live in a one-level-deep reference file, and there are no reference docs at all.

4 / 5

Total

16

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, distinct description with concrete actions and good natural keywords. Its main weakness is the missing explicit "Use when..." trigger clause, which caps completeness and leaves invocation guidance implicit.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks to run device tests for Controls/Core/Essentials/Graphics/BlazorWebView, or to test on an iOS simulator, Android emulator, MacCatalyst, or Windows."

Include the common natural phrases "simulator" and "emulator" among the trigger terms, since users typically say "test on the iOS simulator" or "run tests on an Android emulator".

Mention the available test projects (Controls, Core, Essentials, Graphics, BlazorWebView) so the description covers the full scope of what can be run.

DimensionReasoningScore

Specificity

The description names concrete actions — "Build and run .NET MAUI device tests locally with category filtering" and "Use TestFilter to run specific test categories" — plus explicit platform support. It falls just short of anchor 5 because coverage is limited to build/run/filter without mentioning build-only, rebuild, or class-level filtering.

4 / 5

Completeness

The "what" is clear (build and run MAUI device tests with category filtering, with platform support), but there is no "Use when..." clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is a strong 3 — the platform support detail partially implies usage context — but not anchor 4, which requires both what and when.

3 / 5

Trigger Term Quality

Natural terms users would say are present: "device tests", "iOS", "MacCatalyst", "Android", "Windows", "test categories". A few common variations are missing — "simulator", "emulator", and project names like Controls/Core — so it does not reach comprehensive anchor 5.

4 / 5

Distinctiveness Conflict Risk

" .NET MAUI device tests" is a clear niche with distinct technical triggers (TestFilter, platform/host matrix); it is unlikely to fire for unrelated skills. No overlap risk with general test-running or build skills beyond its own domain.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dotnet/maui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.