CtrlK
BlogDocsLog inGet started
Tessl Logo

scikit-learn

Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.

52

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/coding/scikit-learn/SKILL.md

The canonical home for this skill is scikit-learn in administrakt0r/AI-Agents-Safe-Coding-Skills

SKILL.md
Quality
Evals
Security

Quality

Content

46%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with a genuine one-level-deep reference bundle and clearly signaled navigation, and its workflows are legibly sequenced. But it is noticeably padded with sections that repeat the description, the reference index, and the Quick Start examples; every installation command is broken ("uv uv pip install"); and workflow snippets routinely use undefined variables and missing imports with no validation checkpoints.

Suggestions

Fix the install commands ("uv uv pip install scikit-learn" → "uv pip install scikit-learn") and make Quick Start / Common Workflows snippets self-contained: define X/y (e.g., load a dataset with sklearn.datasets.load_iris) and include the missing OneHotEncoder, numpy, and matplotlib imports.

Cut the padded/redundant sections: drop "When to Use This Skill" (duplicates the description), remove the "Reference Documentation" index that re-lists the inline "See:" links, and delete or merge "Common Workflows" with Quick Start and the Best Practices items that restate textbook knowledge (fit-only-on-train, random_state, stratified splits).

Add validation checkpoints to the workflows: after GridSearchCV report best_score_ and compare against a baseline before predicting on the test set, and in the clustering workflow guard the silhouette-based k selection (e.g., skip k values producing fewer than 2 distinct labels).

DimensionReasoningScore

Conciseness

The ~500-line body has several padded or redundant sections: "When to Use This Skill" repeats the frontmatter description, "Reference Documentation" re-indexes the same six files already pointed to by inline "See:" links, "Common Workflows" re-derives the Quick Start pipeline examples, and "Best Practices" teaches concepts Claude already knows ("Never fit on test data", "Set Random State for Reproducibility", "Fit on Training Data Only"). This matches anchor 2 ("several unnecessary explanations or padded sections") rather than anchor 3's single minor inefficiency; it avoids anchor 1 only because most sections still carry usable code and lists.

2 / 5

Actionability

Most guidance is real, runnable scikit-learn code, but there are systematic execution gaps: every install command is broken ("uv uv pip install scikit-learn" — duplicated 'uv', three occurrences), the Quick Start snippets use undefined variables (X, y, numeric_features, categorical_features), the Common Workflows snippets reference OneHotEncoder, np, and plt without imports, and workflow step 3 builds a ColumnTransformer on variables never defined in that workflow. These are missing key details beyond anchor 4's "minor gaps", landing on anchor 3 ("concrete guidance but incomplete... missing key details") rather than anchor 2, since the code is genuine (not pseudocode) and the bulk is executable once the variables exist.

3 / 5

Workflow Clarity

"Common Workflows" presents clearly numbered, sequenced steps with code for both classification (load → split → preprocess → build → tune → evaluate) and clustering, which is above anchor 2. However, there are no validation checkpoints: nothing verifies the grid search actually improved over the baseline, no check on cv_results_/best_score_ before predicting, and the clustering workflow picks k by argmax with no guard against degenerate silhouette scores — matching anchor 3 ("steps listed but validation gaps; checkpoints missing or implicit"). The Troubleshooting section partially compensates with error recovery (ConvergenceWarning, overfitting, memory), which keeps it from scoring lower, but it is reactive rather than embedded in the workflows.

3 / 5

Progressive Disclosure

The bundle structure is genuinely good: six real reference files (supervised_learning.md, unsupervised_learning.md, model_evaluation.md, preprocessing.md, pipelines_and_composition.md, quick_reference.md) are each signaled inline with "See: references/<file>.md" in the matching capability section, one level deep with no nested chains, and two runnable scripts are documented with what they demonstrate. This sits between anchors 4 and 5: navigation is clear and well-signaled (anchor 5 quality), but the SKILL.md body itself inlines substantial content that belongs in the references — the duplicated Reference Documentation index, full Common Workflows code, and Best Practices material — so organization has minor gaps, giving anchor 4.

4 / 5

Total

12

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description with an explicit, well-scoped "Use when..." trigger clause covering the library's main task areas and good natural keyword coverage. Its main weakness is that the "what" portion names the domain rather than concrete actions, leans on a padded "comprehensive reference documentation" sentence, and omits common synonyms like "sklearn" and "cross-validation".

Suggestions

Lead the "what" clause with concrete actions (e.g., "Trains, evaluates, and tunes classification, regression, and clustering models with scikit-learn") instead of the generic "Machine learning in Python" plus the padded "Provides comprehensive reference documentation" sentence.

Add natural synonyms users actually say — "sklearn" (the import name), "cross-validation", and "train a model/classifier" — to the trigger terms.

Trim the final "comprehensive... best practices" sentence, which duplicates information already conveyed by the task list.

DimensionReasoningScore

Specificity

The description names the domain ("Machine learning in Python with scikit-learn") and lists several concrete task areas — "supervised learning (classification, regression)", "clustering, dimensionality reduction", "model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines". These are task categories rather than the fully concrete actions of the anchor-5 example ("Extract text and tables... fill forms... merge documents"), and the closing sentence "Provides comprehensive reference documentation" is mild padding, so it sits at anchor 4 rather than 5; it is well above anchor 3 because coverage spans the library's whole surface with only minor gaps.

4 / 5

Completeness

Both parts are explicit: the "what" ("Machine learning in Python with scikit-learn... Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices") and a strong explicit "when" ("Use when working with supervised learning..., unsupervised learning..., model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines"). The "when" clause is anchor-5 quality, but the "what" portion only names the domain and the bundle contents rather than concrete actions (weaker than the anchor-5 example's action-listing "what"), placing it between anchors 4 and 5; the score of 4 reflects that the "what" could carry more weight.

4 / 5

Trigger Term Quality

Natural phrases users would actually say are present: "machine learning", "classification", "regression", "clustering", "model evaluation", "hyperparameter tuning", "preprocessing", "building ML pipelines", "scikit-learn". It misses common variations users say — notably "sklearn" (the colloquial import name) and "cross-validation" — matching anchor 4 ("Good keyword coverage; a few natural terms missing") rather than anchor 5's comprehensive synonym/extension coverage, and clearly above anchor 3's partial keyword set.

4 / 5

Distinctiveness Conflict Risk

The scikit-learn/classical-ML niche is clearly staked with distinct triggers ("clustering, dimensionality reduction", "hyperparameter tuning", "ML pipelines"), so it is unlikely to fire for unrelated skills. There is minor overlap risk with closely related skills (deep-learning frameworks like PyTorch/TensorFlow, or general Python data-science skills) since "Machine learning in Python" is broad — matching anchor 4 rather than anchor 5's "clear niche with minimal conflict risk", and well above anchor 3.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (531 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.