CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/risk-matrix-calibration

Checks an already-written risk matrix against what actually happened. Maps each row's likelihood rating to observed defect density, test failure rate and code churn, maps its impact rating to the severity mix and escape rate, then classifies each row as over-stated, under-stated, in-agreement, or not calibrated, using stated reporting thresholds so small differences are not treated as findings. Every proposed rating change carries the observation that produced it, and every proposal is handed to the matrix owner rather than applied. Also emits candidate new entries for areas that show up in defect data but have no row. Owns calibration only: choosing a scoring methodology, designing the matrix structure, picking risk categories, mapping risks to test types, FMEA scoring, review cadence and file storage are all out of scope. Use when a matrix has been driving test decisions for at least three releases and nobody has yet checked whether its ratings match the defects, escapes and incidents that followed.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sequenced, highly actionable calibration procedure with explicit thresholds, validation checkpoints, and feedback loops, supported by two clearly signaled one-level-deep reference files. Its main weakness is verbosity in the conceptual framing, repeated ISTQB definition quotes, and a Limitations section that restates earlier guidance.

Suggestions

Tighten the 'What calibration answers' and 'Before anything else' framing sections, and move the full ISTQB definition quotes (with their URLs) into references/empirical-basis.md, keeping only the term and the operative rule inline.

Collapse the Limitations section so it does not restate points already made in the biased-data section and Step 3.1; cross-reference instead of repeating the 'quiet rows are ambiguous' reasoning.

Trim the 'Scope boundary' prose and the calibration-vs-prediction paragraph, since the description already states the out-of-scope boundary; keep only the table and the one-sentence authoring-skill pointer.

DimensionReasoningScore

Conciseness

The procedure (Steps 1-5) is efficient and specific, but conceptual framing sections ('A risk matrix records what a team believed...'), fully quoted ISTQB definitions with URLs repeated inline, and a Limitations section that restates earlier points ('Quiet rows are ambiguous by construction. This is restated here because...') could be tightened; it is above 1 because the padding is framing around genuinely domain-specific guidance, not explanation of basic concepts Claude lacks.

2 / 3

Actionability

Gives concrete, copy-paste-ready thresholds ('Report at 2 or more points on a 1 to 5 scale', 'fewer than 5 defects as directional only', 'more defects than the median matrix row') and an exact four-part citation requirement (value, source, window, event count); per the instruction-only scoring note, absence of code is not penalized when guidance is this specific, so it clears the 'fully executable...copy-paste ready' anchor over the pseudocode anchor at 2.

3 / 3

Workflow Clarity

Sequenced Steps 1-5 with sub-steps (3.1-3.3), explicit validation checkpoints (Step 3 thresholds gate reporting; Step 4 citation completeness gates reportability), and feedback loops ('If the row has near-zero defects and near-zero test executions...report it under 3.3 instead'); it exceeds the 2 anchor because checkpoints and error-recovery reclassification are explicit rather than implicit.

3 / 3

Progressive Disclosure

Holds the procedure inline while externalizing evidence to references/empirical-basis.md and a full example to references/worked-example.md, both real files clearly signaled with one-level-deep markdown links; it is above 2 because the references are clearly signaled and content that should be separate (evidence, worked example) is split out, not buried inline.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is third-person, specific, and clearly scoped, naming concrete calibration actions and an explicit 'Use when...' trigger that distinguishes it from the authoring skill. It answers both what and when with natural QA terminology and an explicit out-of-scope boundary.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions in third person ('Checks an already-written risk matrix', 'Maps each row's likelihood rating to observed defect density, test failure rate and code churn', 'classifies each row as over-stated, under-stated, in-agreement, or not calibrated', 'emits candidate new entries'), matching the 'lists multiple specific concrete actions' anchor rather than the partial-coverage anchor at 2.

3 / 3

Completeness

Explicitly answers what ('Checks an already-written risk matrix...classifies each row...') and when via an explicit 'Use when a matrix has been driving test decisions for at least three releases and nobody has yet checked...' trigger, satisfying the 'clearly answers both what AND when' anchor; it is above 2 because the when clause is explicit, not implied.

3 / 3

Trigger Term Quality

Covers natural terms a QA/test lead would say ('risk matrix', 'calibration', 'likelihood rating', 'impact rating', 'defects', 'escapes', 'incidents', 'defect density') plus an explicit natural trigger clause, matching the 'good coverage of natural terms' anchor; it is not the 2 anchor because common variations are present rather than missing.

3 / 3

Distinctiveness Conflict Risk

Carves a clear niche with an explicit out-of-scope list ('choosing a scoring methodology, designing the matrix structure, picking risk categories...are all out of scope') and distinct after-the-fact triggers, making it unlikely to trigger for the authoring skill; it is above 2 because the boundary is explicit rather than merely 'somewhat specific'.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents