CtrlK
BlogDocsLog inGet started
Tessl Logo

shap

Model interpretability and explainability using SHAP (SHapley Additive exPlanations). Use this skill when explaining machine learning model predictions, computing feature importance, generating SHAP plots (waterfall, beeswarm, bar, scatter, force, heatmap), debugging models, analyzing model bias or fairness, comparing models, or implementing explainable AI. Works with tree-based models (XGBoost, LightGBM, Random Forest), deep learning (TensorFlow, PyTorch), linear models, and any black-box model.

64

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/coding/shap/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with modern, executable SHAP code, a good explainer-selection decision tree, and well-organized one-level-deep references, but it is significantly over budget: substantial duplication (quick start / workflows / patterns, Key Concepts vs theory.md, Reference Documentation vs the reference files) inflates token cost without adding capability. Trimming to a true overview pointing at the four references would raise both conciseness and progressive disclosure.

Suggestions

Delete the 'Reference Documentation' section — it re-describes the contents of the four reference files, which already exist and are properly indexed in 'Usage Guidelines'.

Move the 'Key Concepts' section (SHAP value interpretation, additivity, background data) into references/theory.md and keep at most the model-output-type warning inline, since it duplicates reference material and Claude's existing knowledge.

Consolidate the duplicated explainer-then-plot code: keep the Quick Start and drop Workflow 1 and Common Pattern 1, or collapse them into a single worked example, and fix the cohort example to use the documented cohort API.

DimensionReasoningScore

Conciseness

At ~570 lines the body is noticeably padded: the 'Key Concepts' section teaches SHAP basics Claude already knows and duplicates references/theory.md; the 'Reference Documentation' section re-describes the contents of all four reference files; and Workflow 1, Common Pattern 1, and the Quick Start repeat the same explainer-then-plot code three times. Not 3: the redundancy is more than 'some unnecessary explanation' — a lean version would cut well over half; not 1: much of the content (decision tree, troubleshooting, performance tips) is genuinely load-bearing.

2 / 5

Actionability

Mostly executable, copy-paste-ready code using the modern API — 'explainer = shap.TreeExplainer(model); shap_values = explainer(X_test)', batching, joblib caching, MLflow logging, and a complete ExplanationService class. Minor gaps keep it from 5: the cohort comparison passes a dict to 'shap.plots.bar' (not the documented cohort API), indexing by feature name ('shap_values[:, "Feature_Name"]') only works with named DataFrame inputs, and the 'Loading references' block is pseudo-guidance formatted as Python comments. Not 3: the code is real and complete, not pseudocode.

4 / 5

Workflow Clarity

The Quick Start decision tree plus six numbered workflows give a clear sequence, and several workflows embed checkpoints ('Validate improvements', 'Check for unexpected feature importance (data leakage)', 'Validate feature relationships make sense'). Minor validation gaps remain — e.g., Workflow 6 (production deployment) lists steps with no verification of the explanation service, and debugging validation is mentioned as a step label rather than an explicit check-and-retry loop. Not 5: checkpoints are named but not operationalized as validate-then-proceed gates; not 3: most workflows do carry explicit validation steps.

4 / 5

Progressive Disclosure

Bundle structure is good: four real reference files (explainers.md, plots.md, workflows.md, theory.md), each one level deep, clearly signaled in the body ('See references/explainers.md ...'), with a 'Usage Guidelines' section telling Claude when to load each. Not 5: content that should live only in the references is also inlined — the Key Concepts section duplicates theory.md and the Reference Documentation section duplicates the reference files' own introductions — which muddies the overview-vs-detail split.

4 / 5

Total

14

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, comprehensive action list, explicit 'Use this skill when...' triggers in third person, and clear SHAP-scoped identity. The only weakness is a few broad trigger verbs (debugging, comparing, bias analysis) that overlap slightly with general ML/XAI skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'explaining machine learning model predictions, computing feature importance, generating SHAP plots (waterfall, beeswarm, bar, scatter, force, heatmap), debugging models, analyzing model bias or fairness, comparing models' — plus the supported model families, which is comprehensive coverage rather than generic language. Not 4: coverage of actions goes beyond 'minor gaps'; every capability the skill body delivers is named.

5 / 5

Completeness

It explicitly answers both questions: 'what' via the concrete action list ('Model interpretability and explainability using SHAP...') and 'when' via the explicit clause 'Use this skill when explaining machine learning model predictions, computing feature importance, ...'. This mirrors the anchor-5 example structure with concrete trigger phrases; not 4 because the 'when' clause is already fully explicit rather than improvable.

5 / 5

Trigger Term Quality

Natural user phrasing is comprehensively covered: 'feature importance', 'explain my model's predictions', 'SHAP', named plot types, 'bias', 'fairness', 'explainable AI', and framework names (XGBoost, LightGBM, TensorFlow, PyTorch) that users would literally say. Not 4: it includes synonyms and the specific tool/framework vocabulary users naturally use, with no meaningful gaps.

5 / 5

Distinctiveness Conflict Risk

The SHAP framing and named frameworks give it a clear niche with minimal conflict risk against unrelated skills, but broad trigger verbs like 'debugging models', 'comparing models', and 'analyzing model bias or fairness' could also fire for non-SHAP debugging or general XAI/interpretability skills. Not 5: those generic triggers carry minor overlap with closely related skills; not 3: the tool-specific scope keeps it well beyond 'somewhat specific'.

4 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (572 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.