Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is highly actionable with a clearly sequenced workflow, cost confirmation before submit, error-recovery guidance, and well-signaled one-level-deep references that verifiably exist in the bundle. Its main weakness is redundancy: the Codex background/heartbeat/session handling is restated in the Workflow, the Command Pattern comments, and the Always-Do-This bullets, which both costs tokens and blurs the file's organization.
Suggestions
State the Codex background/heartbeat procedure once (e.g., in a short 'Runtime-specific download handling' section or a reference file) and reference it from Workflow step 6 and the Command Pattern comments instead of repeating it in three places.
Move the runtime-specific detail (Codex `session_id`/heartbeat polling, Claude Code permission-gating rules) into a small reference file, keeping only the one-sentence rule that applies to every run in the main body.
Group the 'Always Do This' bullets by theme (payload hygiene, paths, runtime/background handling, polling/cost) so related rules read as one block rather than an interleaved grab-bag.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and assumes Claude's competence (no explaining what folding or YAML is), but the Codex background/heartbeat handling is repeated nearly verbatim in three places — Workflow step 6 ("In Codex, run `download-results` as a foreground shell command with `yield_time_ms: 1000`; if Codex returns a `session_id`..."), the Command Pattern comments ("Codex: foreground shell command with yield_time_ms=1000; keep the returned session_id if one is provided"), and two long Always-Do-This bullets. This matches 'Mostly efficient but includes some unnecessary explanation or could be tightened'; the redundancy is more than the 'minor instances' of level 4. | 3 / 5 |
Actionability | The Command Pattern block gives complete, executable `boltz-api` invocations with concrete flags (`--model boltz-2.1`, `--input @yaml:///absolute/path/payload.yaml`, `--raw-output --transform id`), and the payload shapes are shown as real JSON/YAML snippets with exact field names (`chain_ids`, `value`, `binder_chain_id`). Only placeholders the agent must fill (`<job-id-from-start>`, `<run-name>`) remain, and the file explicitly says to replace them — matching 'Fully executable; copy-paste ready code or commands'. | 5 / 5 |
Workflow Clarity | The six-step Workflow is explicitly sequenced (normalize entities → binding block → author payload → `estimate-cost` → confirm → `start` → `download-results`) with a cost-confirmation checkpoint before the paid submit ("run `estimate-cost`, show the USD cost, wait for explicit confirmation") and error-recovery feedback: the 'SAB 400 validation quirk' section says "If the server rejects a payload ... inspect `entities`, `binding`, and `constraints`; read references/results.md for details", and restart guidance covers re-running `download-results` with the same `--name`. This matches 'Clear sequence with explicit validation steps; feedback loops for error recovery'. | 5 / 5 |
Progressive Disclosure | Structure is good: the body is a workflow overview, and both referenced bundle files exist and hold what is claimed — references/api.md carries the payload field reference (verified: entity types, binding variants, bonds, constraints, model_options) and references/results.md carries output layout, nested metrics, and the validation quirk. All references are one level deep and clearly signaled with markdown links. It falls short of level 5 because the runtime-specific operational detail (~25 lines of Codex session_id/heartbeat and Claude Code permission-gating guidance spread across three sections) is inlined in SKILL.md rather than split into a reference with a one-line pointer — 'most content appropriately placed; minor organization gaps'. | 4 / 5 |
Total | 17 / 20 Passed |