Content
85%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is highly actionable — four complete, executable workflow recipes each capped with a validation checklist and an explicit failure-feedback loop for the bundled pipelines. Its one real defect is progressive disclosure: every referenced bundle path (three references/*.md files and three scripts/*.py files) is a dead link because those files were never shipped alongside SKILL.md.
Suggestions
Ship the referenced bundle files — create references/statistical_methods_advanced.md, references/experiment_design_frameworks.md, references/feature_engineering_patterns.md and scripts/experiment_designer.py, scripts/feature_engineering_pipeline.py, scripts/model_evaluation_suite.py — or delete the "Reference Documentation" section and the script-invocation lines; as written, all six referenced paths 404.
Slim SKILL.md to a concise overview (one short signature example per workflow) and move the full function bodies into the reference files, so the ~200-line monolithic body follows the progressive-disclosure structure it already advertises.
Replace the buzzword opener ("World-class senior data scientist skill for production-grade AI/ML/Data systems.") and the generic pytest/black/pylint boilerplate with a one-line statement of what the skill's own workflows do.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dominated by dense, executable code with directive comments rather than explanations of concepts Claude already knows, e.g. the checklists "Never fit transformers on the full dataset — fit on train, transform test". Not 5 because of minor trimmable material: the filler opener "World-class senior data scientist skill for production-grade AI/ML/Data systems." and generic boilerplate commands ("python -m pytest tests/ -v --cov=src/", "python -m black src/ && python -m pylint src/"); not 3 because there is no substantive padding beyond these. | 4 / 5 |
Actionability | All four workflows are complete, copy-paste-ready functions with imports, parameters, and return values (e.g. calculate_sample_size, analyze_experiment, build_feature_pipeline, diff_in_diff), plus concrete CLI invocations with explicit input/output arguments ("--input experiment_spec.json --output experiment_design.json"). This matches the anchor: fully executable code covering the common cases; not 4 because there are no gaps in the code itself. | 5 / 5 |
Workflow Clarity | Each workflow ends with a numbered checklist containing explicit validation checkpoints ("Check for sample ratio mismatch: abs(n_control - n_treatment) / expected < 0.01", "Check overfit_gap > 0.05", "Run a baseline (e.g. DummyClassifier) and verify the model beats it", "Validate parallel trends in pre-period"), and the pipeline section defines an explicit failure feedback loop: "any status other than 'completed' means the stage failed — fix before moving to the next pipeline stage". Not 4 because both checklists for complex processes and error-recovery loops are present, matching the top anchor. | 5 / 5 |
Progressive Disclosure | The "Reference Documentation" section clearly signals three reference files and "Common Commands" invokes three bundled scripts, but none of these exist — there is no references/ or scripts/ directory in the bundle, so all six referenced paths are dead links. Not 4 because missing files entirely is more than "minor organization gaps": navigation is broken and the ~200-line body carries all substantive content monolithically; not 2 because section structure and reference signaling are good, only the targets are absent. | 3 / 5 |
Total | 17 / 20 Passed |