Content
86%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong, high-signal body: copy-paste-ready commands with real outputs, an unusually valuable traps table of verified API pitfalls, explicit validation gates (exit-code gating, incomplete-week flags, run-lag-first precedence), and a clean one-level reference structure that was verified to match the actual bundle. The only slack is illustrative output verbosity and time-stamped example data that is not quarantined from the evergreen guidance.
Suggestions
Trim or abridge the sample output tables (e.g., reporting_lag and mutation_profile) to the rows that teach something, and cut the prose justifying design choices such as the whole-ISO-week widening — the behavior statement alone suffices.
Move date-stamped material ("All verified against the live API on 2026-07-27" and dated example rows) into a short versioned/verification note so evergreen guidance stays separable from time-sensitive claims.
Add one explicit recovery instruction for the name-check gate — e.g., 'if resolve_lineage exits 1, replace withdrawn names with their redesignated successor before continuing' — to close the error-feedback loop the workflow otherwise implies.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and domain-specific — the traps table, per-instance field-name caveats, and the "never from memory" rule are all non-obvious information Claude does not already know, and there is no explanation of generic concepts. It falls at 4 rather than 5 because a few passages are illustrative rather than instructional: full sample output tables (e.g., the reporting_lag output and the four-row mutation table) and prose justifying design decisions ("A window starting mid-week would give a first row covering three days...") could be trimmed, and embedded time-sensitive details ("verified against the live API on 2026-07-27", dated example rows) are not quarantined in a versioned/deprecated section. | 4 / 5 |
Actionability | Every section gives a copy-paste-ready command with real flags: `python3 lineage_prevalence.py --top 5 --where country=USA --weeks 12`, `resolve_lineage.py XFG.23.1.3 PQ.17 PC.2 NOTALINEAGE`, `mutation_profile.py "XFJ*" --versus "XFG*" --gene S --since 2026-01-01`, plus documented output format, `--format table|tsv|json`, and stderr/stdout provenance separation. The examples cover the common cases (discovery, name checking, growth, assay match, lag) — matching the fully-executable 5 anchor. | 5 / 5 |
Workflow Clarity | The sections form a coherent ordered workflow with explicit checkpoints: "Start from the data, not from a remembered list" is labeled the right first command, resolve_lineage's exit code 1 "gates a manuscript's lineage list", incomplete weeks are flagged and excluded unless overridden, and "Run this **before** quoting any recent prevalence" states the lag-check precedence explicitly. It sits at 4 rather than 5 because the ordering lives mostly in section arrangement rather than a stated sequence, and there is no explicit error-recovery loop (e.g., what to do when a name resolves as withdrawn mid-analysis, beyond the PC.2 example). | 4 / 5 |
Progressive Disclosure | SKILL.md is a genuine overview: the four scripts are summarized in a table, full trap detail is deferred with "full detail in `references/lapis-api.md`", and a References section lists all three reference files (lapis-api.md, lineage-nomenclature.md, surveillance-caveats.md — verified present) one level deep, each with a clear scope description. Content is appropriately split and navigation is easy, matching the 5 anchor. | 5 / 5 |
Total | 18 / 20 Passed |