CtrlK
BlogDocsLog inGet started
Tessl Logo

bayesian-inference

Bayesian parameter estimation with MCMC (emcee) and probabilistic programming (PyMC). Posterior distributions, corner plots, model evidence, convergence diagnostics. Use when you need full posterior distributions, not just point estimates.

68

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable: five sequential, executable workflows with explicit convergence thresholds and error-recovery guidance. Its weaknesses are structural — everything lives inline in one long file with no reference files (and no PyMC content despite the description promising it) — plus minor redundancy between the Overview and the description.

Suggestions

Split advanced material (Gelman–Rubin math, model comparison / nested sampling guidance) into one-level-deep reference files (e.g., references/diagnostics.md, references/model-comparison.md) and signal them from SKILL.md.

Either add a PyMC workflow or drop 'probabilistic programming (PyMC)' from the description so the description matches the body's actual coverage.

Trim the Overview section, which restates the frontmatter description almost verbatim, and fold the 'emcee Tips' list into the workflow where each tip applies.

DimensionReasoningScore

Conciseness

Largely code-driven and lean, assuming Claude's competence — no padding explaining what MCMC or Bayesian inference is. Minor over-explanation remains: the Overview nearly duplicates the frontmatter description, and a few code comments restate well-known steps. Not 5 ('every token earns its place') because of that redundancy.

4 / 5

Actionability

Workflows 1–4 are fully executable copy-paste code covering the common cases end-to-end (sampling, corner plots, diagnostics, posterior predictive check), and workflow 5's executable portion runs while the commented dynesty alternative is explicitly justified by the harmonic-mean unreliability warning. Not 4 because there are no real gaps in executable coverage.

5 / 5

Workflow Clarity

A clear sequence (define model → sample → diagnose convergence → posterior predictive check → compare models) with explicit validation thresholds ('Chain length / tau ... Should be > 50', 'R-hat ... Should be < 1.01') and genuine feedback loops: an AutocorrError catch-and-warn and a troubleshooting table mapping symptoms (e.g., 'Trace plots show drift') to fixes ('run longer or improve initialization').

5 / 5

Progressive Disclosure

Well-sectioned single file with no bundle files, but at ~190 lines all content is inlined — the Gelman–Rubin implementation, model-comparison guidance, and troubleshooting could live in one-level-deep reference files, and PyMC is named in the description yet has zero coverage in the body. Matches anchor 3 ('some structure but ... content that should be separate is inline'); not 4 because there are no clearly signaled references at all, and the simple-skill (<50 lines) exception does not apply.

3 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: specific, third-person, and niche-defining with an explicit use-when trigger. The main improvement opportunity is broadening the trigger terms and 'when' clause to cover more natural user phrasings (priors, credible intervals, uncertainty, Bayes factors).

DimensionReasoningScore

Specificity

Names multiple concrete actions — 'parameter estimation with MCMC (emcee)', 'posterior distributions', 'corner plots', 'model evidence', 'convergence diagnostics' — giving comprehensive coverage of the skill's capabilities. Not 4 because the action list spans the full estimation workflow with no meaningful gaps.

5 / 5

Completeness

Clearly answers 'what' (estimation with MCMC/emcee/PyMC, posteriors, corner plots, evidence, diagnostics) and has an explicit 'Use when you need full posterior distributions, not just point estimates' trigger. Not 5 because the 'when' clause covers a single trigger and omits other natural invocation contexts (mentions of priors, MCMC, or uncertainty quantification).

4 / 5

Trigger Term Quality

Good natural-keyword coverage ('Bayesian', 'MCMC', 'posterior distributions', 'point estimates'), but common user phrasings like 'credible intervals', 'priors', 'Bayes factors', or 'uncertainty' are absent. Not 3 because the terms present are natural and specific rather than generic; not 5 because synonym/variation breadth is missing.

4 / 5

Distinctiveness Conflict Risk

'MCMC (emcee)', 'PyMC', and 'Bayesian parameter estimation' carve out a clear niche with distinct triggers that would not fire for generic fitting or plotting skills. Minimal conflict risk with sibling skills like a basic curve-fitting skill.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.