CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/model-risk-evidence-matrix

Assigns an ML model to a low, medium, or high risk tier from what its predictions decide about people, then derives the fairness and explainability evidence that tier must produce: group metrics per declared sensitive feature, intersectional breakdowns with per-cell counts, vulnerability scan categories, a drift monitoring plan, and per-prediction explanation logs. Supplies conventional demographic parity difference bands, a per-vulnerability-category blocking table, evidence rules marking a bundle incomplete or self-contradicting, and a fairness gating workflow that walks a candidate's model card + evidence bundle to a promote / needs-work / block verdict with refuse rules; a reference covers producing the explanation records with Alibi Explain. Use when a model release candidate is up for promotion and someone must decide which fairness artifacts are mandatory, when a declared risk tier's evidence bundle must be checked against what the tier demands, or when the evidence review must gate the promotion.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured decision matrix that adds genuinely non-obvious domain knowledge (regulatory triggers, metric definitions, conflict results, convention bands, evidence rules) and a clear gated promotion workflow with validation. The only slack is minor verbosity in a few quoted framings and a handful of steps that point to external sources rather than inlining the command.

DimensionReasoningScore

Conciseness

Dense with domain-specific knowledge Claude lacks (regulatory citations, exact metric definitions, the fairness-criterion conflict theorem, convention bands, blocking table, evidence rules) with no padding on basics Claude already knows; a few passages (e.g., repeated Fairlearn framing quotes) could be trimmed, so it is efficient but not maximally lean.

4 / 5

Actionability

Provides executable guidance: exact DPD thresholds (0.05/0.10), concrete rule triggers (R1–R6), jq commands, and a full worked output sample; a few steps defer to 'read at the source' rather than giving the inline command, leaving minor gaps.

4 / 5

Workflow Clarity

Steps 1–9 are clearly sequenced, and the Step 9 gating workflow is an explicit ordered checklist with validation checkpoints (re-tier, R1–R6 checks) and refuse-to-promote rules that form a feedback loop, satisfying the validation requirement for a gating/batch operation.

5 / 5

Progressive Disclosure

The body is a well-organized overview that splits detail into two clearly signaled, one-level-deep real references ([references/alibi-explainability.md], [references/worked-example.md]), both verified present, with no nested reference chains.

5 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that names concrete capabilities and gives an explicit, multi-scenario 'Use when' trigger. It is tightly scoped to ML model risk-tiering and fairness evidence gating with low conflict risk; only the trigger phrasings lean slightly technical rather than conversational.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions with comprehensive coverage: 'Assigns an ML model to a low, medium, or high risk tier', 'group metrics per declared sensitive feature, intersectional breakdowns with per-cell counts, vulnerability scan categories, a drift monitoring plan, and per-prediction explanation logs', plus DPD bands, a blocking table, evidence rules, and a gating workflow.

5 / 5

Completeness

Explicitly answers both 'what' (derives the required fairness/explainability evidence and supplies bands, blocking table, rules, and a gating workflow) and 'when' via a concrete 'Use when a model release candidate is up for promotion... or when the evidence review must gate the promotion' clause.

5 / 5

Trigger Term Quality

Good keyword coverage with domain-natural terms ('model card', 'evidence bundle', 'risk tier', 'promote / needs-work / block', 'model release candidate'), but the phrasings are mostly technical and a few natural synonyms a user might say are missing, so it sits below the comprehensive anchor.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (ML fairness/explainability evidence gating for model promotion) with distinct, specific triggers and minimal overlap risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents