Content
88%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body delivers expert, actionable annotation rules in a clear six-step sequence with real validation checkpoints and executable summary code, and it wastes almost no tokens on concepts Claude already knows. The main structural gap is that everything lives in one file; the lineage marker table and subtype rules are natural candidates for a one-level-deep reference file.
Suggestions
Move the lineage/marker table (section 2) and the subtype-bound rules (section 3) to a references/ file (e.g. references/lineage-markers.md), keeping SKILL.md as a leaner overview that links to it — this improves progressive disclosure and trims context load.
Tighten the opening framing paragraph and rhetorical asides (e.g. "They are the clusters that decide a 0.95 agreement gate") into direct statements of the rules they introduce.
Extend the code block slightly to show how the positivity cut is derived from a bimodal split, since the prose offers both a 90th-percentile default and a bimodal alternative but only the percentile version is executable.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and assumes domain competence (no explanation of what MIBI or CD markers are), with every section carrying operational rules. Not level 5 because minor padding could be trimmed: the framing paragraph ("Both directions cost equally under an expert-agreement grader... Everything below follows from that") and rhetorical asides like "They are the clusters that decide a 0.95 agreement gate". Not level 3 because there is no genuine over-explanation of known concepts. | 4 / 5 |
Actionability | Guidance is fully executable: a copy-paste-ready pandas/numpy block computing mean, z-score, and positivity fraction per cluster, a concrete lineage-by-marker table ("CD68, CD163, CD14, CD11b, CD11c, HLA-DR"), and decision rules specific enough to apply directly ("CD68+ CD163-high → M2", "FoxP3+ CD4+ → regulatory"). The remaining steps are judgment calls with explicit criteria and an evidence-file requirement. | 5 / 5 |
Workflow Clarity | Sections 1–6 form a clear sequence (summarise per dataset → lineage → subtype bounds → marker-poor exclusion → sanity checks → two annotators and adjudication), with explicit validation checkpoints in section 5 (class coverage, prevalence plausibility, vocabulary exactness) and an error-recovery loop ("a call you cannot justify from a marker is a call to revisit"). Not level 4 because checkpoints are explicit and batch output is validated before writing. | 5 / 5 |
Progressive Disclosure | The single file is well organized with clear section headers, no nested or dead references, and the code and pitfalls are appropriately placed. Not level 5 because at ~160 lines it is monolithic: the 8-lineage marker table and the subtype rules are reference-grade material that would fit a references/ file, keeping SKILL.md as a leaner overview. Not level 3 because what is inline is navigable and clearly signaled. | 4 / 5 |
Total | 18 / 20 Passed |