Content
85%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-engineered skill body: the data-driven model-selection table makes the hardest decision (which reference model) mechanical, the run step is executable with expected output, and the gotchas checklist encodes real domain pitfalls (scale mismatch, ceiling effects, model shopping). Remaining gaps are small — one duplicated CI interpretation line and only one fully worked command example out of five tools.
Suggestions
State the CI interpretation once (Step 3) and let Step 1 reference it, removing the near-verbatim duplication.
Add short runnable `tu run` examples for the Loewe/CI/ZIP cases (the ZIP dose-matrix one especially), since only Bliss currently has a complete command.
Trim the Step 1 null-model formulas to just the model-choice guidance, since Claude already knows the Bliss/HSA/Loewe definitions.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and assumes competence — no padding about what drugs or dose-response are, and the model-selection logic is delivered via compact tables ('You measured... | Use model | Tool | Input'). It falls short of anchor 5 because of minor duplication and mild over-explanation: the Chou-Talalay CI interpretation ('CI<1 synergy, =1 additive, >1 antagonism') appears nearly verbatim in both Step 1 and Step 3, and the Step 1 null-model theory (E_a + E_b - E_a*E_b) is knowledge Claude largely already has. Not 3 because every section still carries operational value and nothing reads as filler. | 4 / 5 |
Actionability | Concrete and executable: exact tool names with full input parameter listings in the Step 0 table ('effect_a, effect_b, effect_combination (each a fraction 0-1)', 'doses_a, doses_b, viability_matrix'), a copy-paste bash command with a worked example and expected output ('tu run DrugSynergy_calculate_bliss ... -> expected 0.58, bliss_synergy_score 0.12'), and a runnable script with usage. Not 5 because only one of the five tools gets a complete runnable command; the other four are specified by parameter name only, so the common Loewe/CI/ZIP invocations are not copy-paste ready. | 4 / 5 |
Workflow Clarity | A clean numbered sequence (Step 0 pick model by data -> Step 1 understand the null -> Step 2 run -> Step 3 interpret -> Step 4 gotchas) with explicit checkpoints: the scale-conversion gate in Step 0 ('Effects must be on a consistent inhibition scale... convert: inhibition = 1 - viability/100'), the error signal that 'the tools say so' when the Hill fit fails on <3 dose points, and a gotchas checklist covering ceiling effects, model shopping, and score-vs-efficacy. The calculations are non-destructive, so no validate/retry loop is required. Not 4 because validation checkpoints are explicit and woven into the flow rather than merely implied. | 5 / 5 |
Progressive Disclosure | The SKILL.md is a compact overview with well-labeled sections and two clearly signaled, one-level-deep bundle pointers: 'scripts/synergy_reference.py' (verified to exist, and its docstring matches the body's description of it) and a 'Related skills' list routing to sibling skills for curve fitting and pre-computed lookups. No nested references, no content that clearly belongs in a separate file — the inline tables are short enough to justify inlining for a skill this size. Not 4 because navigation is unambiguous and every pointer resolves to a real file with a stated purpose. | 5 / 5 |
Total | 18 / 20 Passed |