CtrlK
BlogDocsLog inGet started
Tessl Logo

vm-lab

Parallels macOS VM lab: GUI automation, Peekaboo, TCC, Ghostty.

64

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/vm-lab/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

90%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary operational skill: dense, executable, niche-specific, and free of fluff, with real validation (two-way signals) and recovery guidance. The only structural improvements are an explicit ordered workflow (or checklist) and moving the pitfall catalog and baseline runbook into reference files.

Suggestions

Add a short numbered 'Typical session' checklist at the top (preflight → clone/discover → run via Ghostty launcher → two-way validate → report) so the section order becomes an explicit sequence — this would lift workflow clarity to 5.

Move the 'Known Pitfalls' catalog and the 'Peekaboo VM Baseline' runbook into a references/ file (e.g., references/pitfalls.md), keeping SKILL.md as a lean overview with clearly signaled links — better progressive disclosure.

State once, near the VM Discovery commands, that 'macOS Tahoe' and the steipete user are placeholders to substitute with the task-owned VM/user confirmed in preflight, so the example commands can't be pasted against the wrong VM.

DimensionReasoningScore

Conciseness

The body is lean and command-dense: every section is terse rules or executable commands ('Only task-owned clones are disposable', 'Use PRL_KEY_ENTER = 36'), with zero padding and no explanation of concepts Claude already knows. Not 4 — there is essentially nothing to trim; every token carries non-obvious, hard-won operational knowledge.

5 / 5

Actionability

Fully executable throughout: copy-paste-ready prlctl/sips commands, a complete heredoc guest launcher script, a bundled typing script (scripts/parallels_type.py, which exists), and specific key-code constants. Example VM names are explicitly flagged as examples in the Bootstrap Preflight section, closing the one gap that would otherwise cost a point.

5 / 5

Workflow Clarity

A clear implicit sequence (Safety Rules → Bootstrap Preflight → VM Discovery → TCC attribution → Ghostty launcher → Two-Way Validation → Baseline → Reporting) with explicit validation checkpoints ('verify through two independent signals', dimension comparison commands) and error-recovery feedback loops ('If keystrokes produce garbage, send Return to clear the line, create a shorter launcher, then retry'). Not 5 because no explicit ordered checklist or step numbering ties the topic sections into a single end-to-end workflow; a reader must infer order from section arrangement.

4 / 5

Progressive Disclosure

Well-organized sections with clearly signaled, one-level-deep references that actually exist ([Bootstrap diagnostics](references/bootstrap-diagnostics.md) and scripts/parallels_type.py are both present in the bundle). Not 5 because nearly all operational detail lives inline in SKILL.md — the Known Pitfalls list and the Peekaboo VM Baseline sequence are candidates for a references/ file, keeping the overview lighter.

4 / 5

Total

18

/

20

Passed

Description

60%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is highly distinctive and uses real niche terminology, but it is a keyword fragment rather than a capability statement: no actions, no 'Use when' trigger guidance, and no statement of what the lab actually does. Adding a verb-led what clause and an explicit 'Use when...' clause would move it to the top band.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when testing GUI automation, TCC permission prompts, or screenshot/click/typing behavior in a clean macOS VM' — this directly lifts the completeness cap.

Replace the noun list with concrete actions, e.g., 'Runs tools inside a Parallels macOS guest and verifies them from the host via screenshots and two-way validation' — this raises specificity from 2 to 4-5.

Include the natural trigger phrases users would say ('screenshot capture', 'clicking and typing', 'VM test run', 'Peekaboo validation') to improve trigger-term coverage.

DimensionReasoningScore

Specificity

The description names the domain concretely ('Parallels macOS VM lab: GUI automation, Peekaboo, TCC, Ghostty') but contains no verbs or concrete actions — it is a noun/keyword list. This matches 'names the domain but actions are minimal or generic' (2); it falls short of 3 because no actual capability (e.g., 'captures screenshots', 'runs guest commands', 'validates GUI actions') is stated.

2 / 5

Completeness

The 'what' is present but only as a fragment — a lab covering these tools — and there is no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3. It is above 2 because the domain and scope are named clearly, not vague.

3 / 5

Trigger Term Quality

Strong natural keywords for the niche: 'Parallels', 'macOS', 'VM', 'GUI automation', 'Peekaboo', 'TCC', 'Ghostty' — terms a user working in this lab would actually say. Not 5 because common phrasings users would use are missing (e.g., 'screenshot capture', 'clicking/typing tests', 'test in a VM', 'two-way validation').

4 / 5

Distinctiveness Conflict Risk

A clear niche with distinct triggers: 'Peekaboo', 'TCC', 'Ghostty', and 'Parallels' are specific tool names unlikely to appear in another skill's remit. Minimal conflict risk.

5 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
steipete/agent-scripts
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.