Content
90%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A dense, accurate CLI-driving skill: fully executable commands, prerequisite self-checks, and strong safety rules for costly and destructive operations. The main gaps are the absence of post-run validation/error-recovery loops and no use of progressive disclosure to move reference material out of the main body.
Suggestions
Add a short post-run verification loop to Steps 5–6, e.g. after bench: 'check the session summary in .skvm/log/bench/<sessionId>/ and surface failures before reporting success', and for detached runs: 'if run-status.json shows failed, show the tailed run.log errors and stop rather than accepting'.
Move the --conditions catalog, proposal-id format details, and the environment-variable table into a single reference file (e.g. references/cli-reference.md) linked from the relevant steps, keeping SKILL.md as a lean overview.
Include an explicit error-recovery path for prerequisite failures beyond exiting (e.g. after the key check, suggest where the user finds their OPENROUTER_API_KEY) so the workflow has a feedback loop rather than a dead stop.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Every section carries skvm-specific facts Claude cannot know — real flag sets, pass semantics, condition strings, proposal id format, env vars — with zero padding and no explanation of general concepts. It matches 'lean and efficient; every token earns its place' rather than level 4, which reserves room for trimmable over-explanation. | 5 / 5 |
Actionability | All seven steps give copy-paste-ready executable bash commands with the real flag set (e.g. 'skvm bench --model=<id> --conditions=original,aot-compiled', 'skvm proposals accept <id> --round=2') plus inline comments stating each example's purpose. Fully executable and covering the common cases, matching the level-5 anchor. | 5 / 5 |
Workflow Clarity | Steps 1–7 are clearly sequenced with pre-flight checks and explicit confirmation gates for expensive/destructive operations ('Never run bench or profile across many models without explicit user confirmation', 'Never run proposals accept unless the user explicitly asked to deploy'). It falls short of level 5 because validation is purely pre-run: there are no post-run verification or error-recovery loops (e.g., checking a bench session's results or handling a failed detached run beyond describing run-status.json). | 4 / 5 |
Progressive Disclosure | The single SKILL.md is well structured with clear per-command sections and a rules section, and no external references are buried or nested. It sits at level 4 rather than 5 because at ~145 lines some self-contained blocks (the env-var table, the --conditions catalog, proposal-id format details) could be split into one-level-deep reference files, and the simple-skill (<50 line) exception does not apply. | 4 / 5 |
Total | 18 / 20 Passed |