Content
90%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A tight, operational skill body: executable code for every common case, environment-specific gotchas Claude could not infer (FlashAttention default, torchtext shim, manifest casing), and a clearly sequenced remote-compute workflow with a strong troubleshooting table. The only weaknesses are the absence of an explicit output-validation step after the batch embed job and a monolithic single-file layout that inlines detail a reference file could hold.
Suggestions
Add a post-job validation checkpoint to the remote-compute workflow, e.g. after save_artifacts verify `embedded.obsm['X_scGPT'].shape[0] == adata.n_obs` before declaring success.
Consider moving the Troubleshooting table and remote-compute handle details into a reference file (e.g. references/troubleshooting.md) to keep SKILL.md a lean overview, since the body currently inlines all detail.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is lean and operational throughout — no explaining what AnnData, foundation models, or clustering are; it jumps straight to loading vocab, calling embed_data with annotated parameters, and a compact prerequisites table. The one longer narrative passage (remote-compute handle semantics: "`.close()` lives on the handle, not on the job object") is harness-specific knowledge Claude would not know, so it earns its tokens. Every section maps to the score-5 anchor ("assumes Claude's competence; every token earns its place"); it is not score 4 because there is no identifiable padding to trim. | 5 / 5 |
Actionability | Core flows are copy-paste ready: GeneVocab.from_file with a verification print, a complete embed_data call with typed arguments, a full submit_job invocation, and the attach/close sequence. The few placeholders ("/path/to/scgpt-human", "environment=...") are explicitly justified — each is annotated with where to obtain the real value ("env name from compute_details", checkpoint path from compute_details), and troubleshooting gives an exact fix command (`name.lower().replace('-', '_')`). This matches the score-5 anchor ("specific examples cover the common cases") rather than 4's "minor gaps", since no example is pseudocode or missing key details. | 5 / 5 |
Workflow Clarity | The remote-compute workflow is clearly sequenced: compute_details → host.compute.create → submit_job → wait_for_notification → save_artifacts → attach_job for the full result, with checkpoints like `print(len(gv)) # 60697` and a symptom→fix troubleshooting table (e.g. "Nearly all genes dropped → Wrong gene_col") serving as feedback loops. It falls short of 5 because the batch embed job has no explicit output-validation step (e.g. verifying embedded.h5ad retains the input cell count) — a minor validation gap per the score-4 anchor, not the missing-validation case of 3, since failure symptoms and recovery paths are documented. | 4 / 5 |
Progressive Disclosure | Sections are well-organized (Prerequisites, How to run, Output format, Remote compute, Gotchas, Troubleshooting, Next) and deep orchestration detail is correctly deferred one level to sibling skills ("See the `remote-compute-ssh` / `remote-compute-modal` skill"), with no nested references. It does not reach 5 because the ~120-line body inlines everything (troubleshooting and remote-compute detail could live in reference files), exceeding the under-50-lines case the rubric exempts; structure is good with minor organization gaps, matching the score-4 anchor. | 4 / 5 |
Total | 18 / 20 Passed |