Content
92%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured instruction skill: fully sequenced workflow with explicit fault-tolerance checkpoints, copy-paste-ready R code with honest placeholder handling, and clean one-level-deep references verified on disk. The only weakness is minor redundancy (repeated summary lines, duplicated reference pointers, and overlapping rule sections) that could be trimmed for token efficiency.
Suggestions
Remove the redundant trailing 'Reference Files' table (the Step 4 inline '→' pointers already link both files), or keep only one of the two placements.
Drop the opening restatement of the frontmatter summary ('Always outputs four workload configurations and a recommended primary plan') since Step 2 states the same requirement in full.
Merge the 'Fault tolerance guidelines' into the relevant Hard Rules (6–8 overlap heavily) to eliminate the double-stated weak-instrument guidance.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and assumes domain competence (methods are named with thresholds rather than tutored), but has minor redundancy: the opening line repeats the frontmatter summary ('Always outputs four workload configurations and a recommended primary plan'), the trailing 'Reference Files' table duplicates the inline Step 4 pointers already given, and fault-tolerance rules overlap Hard Rules 6–8. Anchor 4 ('efficient; minor instances... that could be trimmed') fits better than 5's 'every token earns its place'. | 4 / 5 |
Actionability | Guidance is fully executable: concrete thresholds ('p < 5×10⁻⁸, LD clumping r² < 0.001 / 10,000 kb', 'F > 10', 'PP.H4 > 0.8'), a complete copy-paste-ready TwoSampleMR R template with the available_outcomes() lookup pattern, and specific fallback methods ('LIML, sisVIVE'). The EXAMPLE-ID placeholders are explicitly justified with inline comments, and per the scoring notes the instruction-style planning guidance is equally concrete. | 5 / 5 |
Workflow Clarity | Steps 1–9 are clearly sequenced with explicit validation checkpoints and feedback loops: 'If IV count falls below 3: warn the user that MR is not feasible', 'If F-statistic < 10 for all IVs: do not proceed with IVW as primary', 'Confirm this fits within any stated time constraints before recommending', and a 'Revision strategy if first-pass findings fail'. This matches the anchor-5 pattern of explicit validation plus error-recovery loops. | 5 / 5 |
Progressive Disclosure | Both referenced files (references/iv_benchmarks.md, references/gwas_databases.md) exist, contain exactly the promised lookup tables, and are one level deep. They are clearly signaled inline in Step 4 with '→' arrows and summarized in a 'Reference Files' table with a 'Used In' step column, while bulky per-exposure-class data is appropriately split out of the overview. | 5 / 5 |
Total | 19 / 20 Passed |