Content
76%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A lean, highly actionable reference with executable code, useful model/VRAM tables, and a well-sequenced remote-compute workflow. Its main gaps are the unshown score_evo2.py job script and the absence of an explicit result-verification checkpoint in the batch scoring workflow, which caps workflow clarity.
Suggestions
Add a verification checkpoint after the compute_done notification (e.g., inspect scores.json for expected length/range before calling save_artifacts) so the batch scoring workflow has an explicit validation step.
Include the contents of score_evo2.py (or a skeleton showing the HF_HOME/HF_HUB_OFFLINE setup and score_sequences call) so the remote-compute example is fully copy-paste ready.
Move the troubleshooting and typical-performance tables into a reference file to keep SKILL.md as a tighter overview, since no bundle files currently exist.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense tables and executable code with zero explanations of concepts Claude already knows; even prose lines carry non-obvious operational facts ('set HF_HUB_OFFLINE=1 so the loader doesn't try to write refs/ into a read-only mount', 'More negative ⇒ less likely under the model'). Every token earns its place, matching anchor 5 rather than the minor-trimming of anchor 4. | 5 / 5 |
Actionability | Most guidance is executable — 'pip install evo2', 'Evo2("evo2_7b")', 'model.score_sequences(seqs)', 'model.generate(prompt_seqs=[...])', and a concrete submit_job call — but 'score_evo2.py' is passed as a job input without its contents being shown, and 'env selection is host-specific' leaves a gap. This fits anchor 4 (mostly executable with minor gaps), not 5, because the pivotal scoring script is referenced rather than provided. | 4 / 5 |
Workflow Clarity | The remote-compute sequence is clearly ordered (read compute_details → create provider → submit_job → wait_for_notification → save_artifacts → attach_job) and the troubleshooting table provides error recovery, but this is a batch workflow ('score_sequences, 200×200bp') with no explicit verification of job results before acting on the payload. Per the rubric's cap on batch operations lacking validation, this sits at anchor 3, not 4. | 3 / 5 |
Progressive Disclosure | No bundle files exist; the single SKILL.md is well-sectioned (Prerequisites, How to run, Models, Output format, Decision tree, Remote compute, Performance, Troubleshooting) and defers orchestration detail to clearly-named one-level-deep external skills ('remote-compute-ssh' / 'remote-compute-modal'). At ~125 lines, some content (troubleshooting and performance tables) could live in reference files, fitting anchor 4 rather than the clean split of anchor 5. | 4 / 5 |
Total | 16 / 20 Passed |