Content
86%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exemplary lean, dense guidance skill: it teaches only project-specific invariants (per-iteration token tax, write-echo discipline, prompt placement) with named files, commands, and markers. The residual gaps are the absence of an explicit ordered workflow and fully executable inline examples for the benchmarking step.
Suggestions
Add a short ordered checklist at the top (benchmark before → make the change → benchmark after, same models/cases → rerun `python system_prompts/generate.py` if prompts changed) so the implicit workflow becomes an explicit sequence.
Include a minimal good-vs-bad example of a write tool's return shape to anchor the `{ success, message }` invariant concretely.
Name the exact command or entry point in the `ai-evals` skill for running the affected mode, so the benchmark step is copy-paste executable without loading the other skill first.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Every line carries non-obvious, project-specific knowledge — e.g. 'the system prompt **plus every tool schema** is re-sent on every loop iteration', 'Optimize **`finalContextTokens`** (window occupancy — what drives overflow and compaction)', and the canonical write-echo mistake. There is no padding and nothing Claude already knows is re-explained. | 5 / 5 |
Actionability | Highly concrete guidance throughout: named files and helpers ('`finishAppDraftWrite` in `global/core.ts`', '`system_prompts/base/`', '`system_prompts/README.md`'), an exact command ('rerun `python system_prompts/generate.py`'), exact markers ('`<!-- chat-only -->` / `<!-- cli-only -->`'), and an exact return shape ('return `{ success, message }` instead'). It stops short of a 5 because the benchmark step's execution is fully delegated to the external `ai-evals` skill and no code/schema snippets are shown inline. | 4 / 5 |
Workflow Clarity | A clear validation checkpoint is stated explicitly — 'Run the affected mode **before** your change and **after**, same model(s), same cases' — plus a re-check loop for shared refactors ('re-check this invariant for **all** the write tools routing through it'). However, the body is organized as principles under topical headings rather than one ordered sequence, so sequencing must be assembled by the reader rather than followed as a checklist. | 4 / 5 |
Progressive Disclosure | The body is under 50 lines with no bundle files, and content is well organized under six clear section headers. The only external pointers ('see the `ai-evals` skill', '`system_prompts/README.md` has the mechanics') are one level deep and clearly signaled, so nothing needs splitting out. | 5 / 5 |
Total | 18 / 20 Passed |