CtrlK
BlogDocsLog inGet started
Tessl Logo

ui-tests-local-vm

Run PowerToys UITest.Next suites in persistent Hyper-V VMs over PowerShell Direct. Validate full default and constrained suites on Windows 10 and Windows 11, plus applicable ARM64 guests, then automatically hand implementation tasks to ui-tests-pipeline-ci for commit/push and CI validation. Local green is not end-to-end completion. Use for agentic UI-test iteration, reusable standard-user desktops, payload staging, evidence export, VM customization, clean baselines, or hosts without nested virtualization. Keywords: Hyper-V, local VM, virtual machine, PowerShell Direct, Copy-VMFile, VMBus, checkpoint, unattend, autounattend, ISO, Windows 10 LTSC, Windows 11, ARM64, Windows on ARM, UI tests, UITest.Next, winappcli, TRX, CI handoff.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary operational skill: executable commands with full parameter sets, a validated multi-stage workflow with checkpoints and feedback loops, and clean one-level-deep reference navigation. The only real cost is token redundancy — the core gate rules are repeated across multiple sections and could be consolidated without losing compliance emphasis.

Suggestions

Consolidate the Windows-10-then-Windows-11 full-suite gate, currently stated in the intro, Guest OS policy, the mermaid cycle, task-list steps 10-11, and Non-negotiable rules: state it fully once and cross-reference from the other sections.

Merge the repeated 'never autopilot host setup / IsReady=false means BLOCKED' guidance (task-list step 0, Quick start, and Non-negotiable rules) into one authoritative section, keeping one-line pointers elsewhere.

Trim restatements of 'local green is not end-to-end completion' to the intro boundary statement plus the task-list step where the CI handoff actually happens.

DimensionReasoningScore

Conciseness

The body is dense and almost entirely project-specific (NTFS vs Dev Drive vmms wedging, DPAPI credential handling, LTSC build requirements) rather than explaining concepts Claude already knows, and all code blocks are commands, not exposition. But the Windows-10-then-Windows-11 full-suite gate and the 'local green is not completion / never autopilot host setup' rules are each restated across four to five sections (intro, Guest OS policy, the cycle, the task list, Non-negotiable rules), which is trim-able redundancy — more than a single minor instance, so not 5; the rest is efficient, so not 3.

4 / 5

Actionability

Quotes: "pwsh .github\skills\ui-tests-local-vm\scripts\Initialize-LocalVmHost.ps1 -VmRoot X:\PowerToysUiTestVm -CheckOnly" and the full "Invoke-LocalVmUiTest.ps1" invocation with every parameter ("-Filter 'Name=MyModule.FocusedTest' -Platform x64Win10 -BuildLabel (git rev-parse HEAD) -SuiteTimeout 15m"), plus concrete practices like "read the controller result's `.Failed` array (non-passed tests + first error line) instead of re-parsing TRX". Copy-paste-ready commands cover the common cases, matching the top anchor.

5 / 5

Workflow Clarity

The mermaid flowchart sequences design→build→package→run→evidence→widen→both-OS gate→constrained→CI with explicit decision nodes, and the 16-step checklist has real validation checkpoints: "Verify host setup FIRST ... If it reports IsReady=false, STOP", "Build product and test projects on the host to exit code 0", "Run the controller with -PlanOnly and inspect its request/plan", "Probe the non-admin interactive desktop before test execution", and "Always parse TRX and require `total > 0` plus `executed == total`". Feedback loops (focused test → diagnose first controlling failure → rebuild with -ReuseStagedPayload) are explicit, matching the top anchor exactly.

5 / 5

Progressive Disclosure

"Required reads" maps each of the five real references/ files to a one-line scope with read conditions ("**read for any shell-extension module**", "Read only what the task needs"), all verified present; links are one level deep with correct section anchors (references/setup.md#0-... and #3-get-windows-media match actual setup.md headings), and scaffolded scripts (New-UiTestVm.ps1, Reset-LocalVm.ps1) are real targets documented in setup.md. The SKILL.md keeps only overview, policy, and the cycle while details live in the references — the top anchor's structure.

5 / 5

Total

19

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense, third-person description that answers what, when, and for whom with concrete actions, an explicit 'Use for' clause, and a synonym-rich keyword list. Its only weakness is trigger-term overlap with its sibling UI-test skills (migration and pipeline-CI), which share the 'UI tests'/'CI' vocabulary.

DimensionReasoningScore

Specificity

Quotes: "Run PowerToys UITest.Next suites in persistent Hyper-V VMs over PowerShell Direct", "Validate full default and constrained suites on Windows 10 and Windows 11, plus applicable ARM64 guests", "automatically hand implementation tasks to ui-tests-pipeline-ci for commit/push and CI validation". Multiple specific concrete actions with comprehensive coverage of the skill's scope; not below 4 because coverage extends beyond the core run action to validation profiles, guest OSes, and the CI handoff, matching the 'comprehensive coverage' anchor exactly.

5 / 5

Completeness

The 'what' is explicit ("Run ... suites in persistent Hyper-V VMs over PowerShell Direct", "Validate full default and constrained suites", "hand implementation tasks to ui-tests-pipeline-ci for commit/push and CI validation") and the 'when' is explicit with concrete triggers ("Use for agentic UI-test iteration, reusable standard-user desktops, payload staging, evidence export, VM customization, clean baselines, or hosts without nested virtualization"). It also states an explicit boundary ("Local green is not end-to-end completion"), matching the top anchor for both what and when.

5 / 5

Trigger Term Quality

Quotes: "Use for agentic UI-test iteration, reusable standard-user desktops, payload staging, evidence export, VM customization, clean baselines, or hosts without nested virtualization" plus "Keywords: Hyper-V, local VM, virtual machine, PowerShell Direct, Copy-VMFile, VMBus, checkpoint, unattend, autounattend, ISO, Windows 10 LTSC, Windows 11, ARM64, Windows on Arm, UI tests, UITest.Next, winappcli, TRX, CI handoff". Comprehensive natural-term coverage including synonyms ("local VM"/"virtual machine", "ARM64"/"Windows on Arm") and concrete technologies a user with this need would actually say.

5 / 5

Distinctiveness Conflict Risk

Quotes: "persistent Hyper-V VMs over PowerShell Direct", "hosts without nested virtualization", "CI handoff", "UI tests, UITest.Next". The Hyper-V/PowerShell-Direct niche is clear and largely non-conflicting, but trigger terms like "UI tests", "UITest.Next", and "CI handoff" overlap with the closely related sibling skills ui-tests-migration and ui-tests-pipeline-ci mentioned in the description itself — minor overlap risk with closely related skills, so not 5; well above the 'somewhat specific' midpoint, so not 3.

4 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 3 suspicious

Warning

Total

15

/

16

Passed

Repository
microsoft/PowerToys
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.