CtrlK
BlogDocsLog inGet started
Tessl Logo

shap

Model interpretability and explainability using SHAP (SHapley Additive exPlanations). Feature importance, dependence plots, interaction effects, and fairness analysis for any black-box model.

53

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Data Analysis/shap/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well structured and highly actionable for the quick-start path, with clearly sequenced workflows. Its two real problems are significant redundancy across the Reference Documentation / Usage Guidelines / Best Practices sections, and a broken progressive-disclosure chain: all four referenced files are cited but missing from the bundle.

Suggestions

Ship the four referenced files (references/explainers.md, plots.md, workflows.md, theory.md) or remove/deduplicate the pointers — currently every "See references/..." link is dead.

Cut the "Reference Documentation" and "Usage Guidelines" sections down to one-line pointers per file; they restate content already signaled inline, which is the main source of verbosity.

Trim "Best Practices Summary" to points not already covered in Performance Optimization and Key Concepts, and add an explicit feedback loop for the debugging workflow (what to do when a validation check fails).

DimensionReasoningScore

Conciseness

The ~560-line body has several padded/duplicated sections: "Reference Documentation" re-describes the contents of all four reference files in detail (repeating the inline "See references/..." pointers), "Usage Guidelines" restates the same loading guidance, and "Best Practices Summary" repeats points already made in Performance Optimization and Key Concepts. This matches anchor 2 ("Noticeably verbose; several unnecessary explanations or padded sections") rather than 3, because the duplication is structural, not just occasional over-explanation; it is above 1 because the core quick-start and patterns sections are substantive rather than explaining basics Claude already knows.

2 / 5

Actionability

Most guidance is executable: the explainer decision tree, `explainer = shap.TreeExplainer(model)` / `shap_values = explainer(X_test)`, concrete `shap.plots.beeswarm/waterfall/scatter` calls, MLflow logging, joblib caching, and batching code are all copy-paste ready. This matches anchor 4 ("Mostly executable guidance... with minor gaps") — some snippets rely on placeholders like "Most_Important_Feature"/"Suspicious_Feature" and the cohort bar plot passes a dict that may need adaptation — so it does not reach 5.

4 / 5

Workflow Clarity

All six workflows are clearly sequenced with stated goals and numbered steps, and several include validation checkpoints ("Validate feature relationships make sense", "Check for unexpected feature importance (data leakage)", "Validate improvements", "Monitor explanation quality"), matching anchor 4 ("Clear sequence with most checkpoints present; minor validation gaps"). It falls short of 5 because there are no explicit feedback loops (e.g., what to do when validation fails) and workflows 4-6 defer their substance to the non-existent references/workflows.md.

4 / 5

Progressive Disclosure

References are well signaled and one level deep ("See `references/explainers.md`", "See `references/plots.md`", plus a Usage Guidelines section on when to load each), but none of the four cited files (explainers.md, plots.md, workflows.md, theory.md) exist in the bundle — there is no references/ directory at all. Scored against the actual bundle structure, this matches anchor 3 ("Some structure but could be better organized") because the navigation promise is broken and workflows 2-6 lose their detailed content; it is above 2 because the in-body structure itself is genuinely well organized with inline pointers, and below 4 because following any pointer would fail.

3 / 5

Total

13

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, uses natural domain keywords, and is clearly distinguishable as the SHAP skill. Its main weakness is the absence of any explicit 'when to use' trigger guidance, which caps completeness and slightly weakens trigger quality.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks to explain model predictions, compute or visualize SHAP values, check feature importance, or analyze model fairness/bias."

Include a couple of common user phrasings as trigger synonyms ("why did my model make this prediction", "explainable AI", "model debugging") to strengthen trigger term coverage.

Soften "for any black-box model" or add the supported model families (tree, deep learning, linear) to avoid over-claiming and improve specificity.

DimensionReasoningScore

Specificity

Quotes: "Feature importance, dependence plots, interaction effects, and fairness analysis for any black-box model" — four concrete capability areas plus a named method (SHAP). This matches the anchor "Lists several specific actions; minor gaps in coverage" rather than 5, because coverage omits debugging/production use cases and "any black-box model" slightly over-claims; it sits above 3 because it goes well beyond naming the domain with 1-2 actions.

4 / 5

Completeness

The "what" is clear and concrete (SHAP-based interpretability, importance, plots, interactions, fairness), but there is no "Use when..." clause or equivalent explicit trigger guidance anywhere in the description. Per the judging guidelines this caps completeness at 3, exactly matching the anchor "Has a clear 'what' but 'when' is missing or only weakly implied"; it is above 2 because the what is substantive, not vague.

3 / 5

Trigger Term Quality

Quotes: "Model interpretability and explainability", "SHAP (SHapley Additive exPlanations)", "Feature importance", "dependence plots", "interaction effects", "fairness analysis", "black-box model". These are natural terms users would say (matching anchor 4: "Good keyword coverage; a few natural terms missing"), but common phrasings like "explain my model's predictions", "why did my model predict X", or "model debugging" are absent, so it does not reach the comprehensive synonym coverage of 5.

4 / 5

Distinctiveness Conflict Risk

Quotes: "using SHAP (SHapley Additive exPlanations)" — the named library and method give it a clear niche distinct from generic ML skills. It fits anchor 4 ("Mostly distinct; minor overlap risk") rather than 5 because "model interpretability and explainability" broadly could overlap with a LIME/general-XAI or fairness-analysis skill that shares those trigger phrases.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (582 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 11 missing

Warning

Total

13

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.