Content
87%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exceptionally lean, high-value body: executable commands, exact batch CSV schema, and hard-won operational gotchas (silent precompute, RAM ceiling, YAML-over-CLI override, flag-name pitfall) with nothing wasted. The one gap is the absence of an output-validation checkpoint for batch runs, which the rubric caps workflow clarity at 3 for.
Suggestions
Add a short verification step after running (e.g., check that rank1.sdf exists for every complex_name in the batch and sanity-check that the reported confidence is a logit only comparable within the same complex) — this would lift workflow clarity above the batch-operation validation cap.
In the batch path, mention what an empty/failed row looks like in the output directory so a long library run can be checked mid-flight instead of only after completion.
Consider stating in the batch section how many samples per complex the default YAML produces, since the YAML-overrides-CLI gotcha makes the default sampling depth easy to get wrong silently.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Every section carries non-obvious, DiffDock-specific information (YAML config silently overwriting CLI flags, the ~11-minute silent precompute and 32 GB RAM requirement, the '--ligand' argparse prefix-match accident) with zero filler and no explanation of concepts Claude already knows. Lean and efficient — matches the 5 anchor. | 5 / 5 |
Actionability | Fully executable guidance: a copy-paste-ready inference command with real flag values and an example SMILES, the exact four-column batch CSV schema, the output naming convention ('rank{N}_confidence{score}.sdf'), concrete fixes ('sed the constant in inference.py to min(64000, rlimit[1])', 'provider_params.modal.memory: 65536'), and an error-recognition table. Covers the common single-complex and batch cases. | 5 / 5 |
Workflow Clarity | The run → output-interpretation → error-recovery flow is clear and the error table gives symptom-to-fix mapping, but the skill supports batch operations ('--protein_ligand_csv batch.csv', fragment-library screening) with no output-validation or verification checkpoint, and the rubric explicitly caps batch workflows without validation at 3. | 3 / 5 |
Progressive Disclosure | The body is a lean overview that handles the single-complex path inline and clearly signals the one-level-deep split: 'that path and a larger-library screening recipe are in references/workflows.md' — a real, relevant file (verified). Content is appropriately divided with easy navigation. | 5 / 5 |
Total | 18 / 20 Passed |