Content
85%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is highly actionable: five sequential, executable workflows with explicit convergence thresholds and error-recovery guidance. Its weaknesses are structural — everything lives inline in one long file with no reference files (and no PyMC content despite the description promising it) — plus minor redundancy between the Overview and the description.
Suggestions
Split advanced material (Gelman–Rubin math, model comparison / nested sampling guidance) into one-level-deep reference files (e.g., references/diagnostics.md, references/model-comparison.md) and signal them from SKILL.md.
Either add a PyMC workflow or drop 'probabilistic programming (PyMC)' from the description so the description matches the body's actual coverage.
Trim the Overview section, which restates the frontmatter description almost verbatim, and fold the 'emcee Tips' list into the workflow where each tip applies.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Largely code-driven and lean, assuming Claude's competence — no padding explaining what MCMC or Bayesian inference is. Minor over-explanation remains: the Overview nearly duplicates the frontmatter description, and a few code comments restate well-known steps. Not 5 ('every token earns its place') because of that redundancy. | 4 / 5 |
Actionability | Workflows 1–4 are fully executable copy-paste code covering the common cases end-to-end (sampling, corner plots, diagnostics, posterior predictive check), and workflow 5's executable portion runs while the commented dynesty alternative is explicitly justified by the harmonic-mean unreliability warning. Not 4 because there are no real gaps in executable coverage. | 5 / 5 |
Workflow Clarity | A clear sequence (define model → sample → diagnose convergence → posterior predictive check → compare models) with explicit validation thresholds ('Chain length / tau ... Should be > 50', 'R-hat ... Should be < 1.01') and genuine feedback loops: an AutocorrError catch-and-warn and a troubleshooting table mapping symptoms (e.g., 'Trace plots show drift') to fixes ('run longer or improve initialization'). | 5 / 5 |
Progressive Disclosure | Well-sectioned single file with no bundle files, but at ~190 lines all content is inlined — the Gelman–Rubin implementation, model-comparison guidance, and troubleshooting could live in one-level-deep reference files, and PyMC is named in the description yet has zero coverage in the body. Matches anchor 3 ('some structure but ... content that should be separate is inline'); not 4 because there are no clearly signaled references at all, and the simple-skill (<50 lines) exception does not apply. | 3 / 5 |
Total | 17 / 20 Passed |