CtrlK
BlogDocsLog inGet started
Tessl Logo

xgboost-analysis

Use when building XGBoost models on tabular data and returning feature importance ranking outputs. Supports binary classification and regression with automatic task detection, train-test split, performance tables, feature importance ranking tables, and PNG importance plots.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Well-structured, executable content with excellent workflow validation and error-recovery guidance. Main flaws: the tests/data files referenced by the examples and validation commands are missing from the bundle, and parts of the CLI argument table belong in cli-guide.md.

Suggestions

Bundle the tests/data sample files (dt_sample1.csv, dt_sample2.csv, dt_sample3.txt) referenced by the Quick Examples and Validation sections, or rewrite those commands to reference files that actually ship with the skill.

Move the full --core-arguments table to references/cli-guide.md and keep only the required/common flags in SKILL.md to reduce token load on every invocation.

Trim the Feature Importance Metrics explanations to a one-line recommendation (default gain) and defer metric interpretation details to references/algorithm.md.

DimensionReasoningScore

Conciseness

The body is efficient — no concept padding, clear Use/Do-Not-Use sections, table-driven argument docs — but the 23-row core-arguments table and the gain/cover/frequency explanations partly duplicate material that belongs in references/cli-guide.md. It is not 5 because some trimming/moving is possible; not 3 because there is no real over-explanation of things Claude already knows.

4 / 5

Actionability

Concrete, executable commands throughout (primary command with placeholders, package install one-liner, three worked examples, validation run with expected output paths). It is not 5 because the example and validation commands reference tests/data/dt_sample*.csv files that are not present in the bundle, so they fail as written; not 3 because the primary workflow runs fine with the user's own data.

4 / 5

Workflow Clarity

A clear 3-step Minimal Workflow, an explicit Validation section with exact expected output files to verify, documented exit behavior ("exits with SKILL_MISSING_INPUT"), and a Common Errors table with a recovery pointer to troubleshooting — explicit checkpoints and an error-recovery loop. It is not 4 because validation checkpoints are fully explicit, not mostly present.

5 / 5

Progressive Disclosure

Good structure with a "Read These Files When Needed" table mapping needs to real, one-level-deep reference files (algorithm.md, cli-guide.md, troubleshooting.md all exist). It is not 5 because the body advertises "Bundled test data" at tests/data/, which does not exist in the actual bundle — a broken navigation pointer.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit "Use when" trigger, concrete third-person capability list, and a distinct niche. The only weakness is trigger term coverage missing common synonyms and file extensions.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions spanning the full pipeline — "building XGBoost models on tabular data", "automatic task detection, train-test split, performance tables, feature importance ranking tables, and PNG importance plots" — with comprehensive coverage and correct third-person voice ("Supports..."). It is not 4 because there are no meaningful gaps in capability coverage.

5 / 5

Completeness

It explicitly answers both questions: an explicit trigger clause ("Use when building XGBoost models on tabular data and returning feature importance ranking outputs") and a concrete what ("Supports binary classification and regression with automatic task detection, train-test split, performance tables..."). It is not 4 because the "when" clause is explicit and specific rather than merely adequate.

5 / 5

Trigger Term Quality

Natural user terms like "XGBoost", "tabular data", "feature importance", "classification", and "regression" are present, but coverage lacks synonyms (e.g., "gradient boosting") and file extensions (e.g., .csv). It is not 5 because the rubric's top anchor expects comprehensive coverage including synonyms and extensions.

4 / 5

Distinctiveness Conflict Risk

The XGBoost-on-tabular-data-with-feature-importance niche is distinct and unlikely to fire for unrelated skills; trigger terms are specific. It is not 4 because the triggers are clearly distinct with minimal overlap risk, not just "mostly distinct".

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.