CtrlK
BlogDocsLog inGet started
Tessl Logo

severity

Rates how severe a finding or bug is if it is real — the blast-radius axis, complementary to confidence's is-it-real axis. Emits a lowercase tier (critical / high / medium / low) from a fixed axis rubric, then applies one deterministic, executable path floor (auth / billing / migration / infra / secrets paths) plus heuristic escalators (data-loss, security, concurrency shapes) — both scoped to reachable production code, never test / fixture / generated paths. Callers own how they gate on the tier and map it to their own blocking rules; the skill stays policy-free. Use to triage a review finding or a bug. Triggers on "how severe", "rate severity", "severity check", "triage this finding", "how bad is this", "/severity".

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable instruction skill: the tiering rubric is concrete and executable (real case-matching code, explicit ordering, exact output token), and the three-step workflow is well-gated with evidence citation and regression-tested evals. The main weaknesses are moderate repetition of the severity-vs-confidence/policy-free framing and, most importantly, progressive disclosure: the body leans on multiple bundle and repo paths that are not actually present alongside the SKILL.md, leaving its references unresolvable and its consumer-integration detail inlined.

Suggestions

Ship the referenced files (at minimum `scripts/eval/golden/severity-tiering.jsonl` and the eval harness, plus the consumer-integration docs like conventional-comments.md / review-config.md) inside the skill's references/ or scripts/ directories, or rewrite those mentions as self-contained descriptions so no reference dangles.

Move the repo-specific consumer material (the 'Mapping to a reviewer's blocking flag' crosswalk and 'How Callers Consume the Tier' sections) into a reference file linked one level deep, keeping SKILL.md as the rubric + output format core.

State the severity-vs-confidence orthogonality and the policy-free stance once (the opening blockquote) and drop the re-statements in 'How Callers Consume the Tier' to tighten token efficiency.

DimensionReasoningScore

Conciseness

The core sections (tier table, axes, Step 1–3, output format) are dense and earn their tokens — executable case blocks, tables, and an exact format template with no explaining of concepts Claude already knows. It is not level 5 because points repeat: the severity-vs-confidence distinction appears in the opening blockquote and again in "How Callers Consume the Tier" ("it is advisory, like confidence, and it is policy-free"), and the policy-free point is made twice. Not level 3 because the padding is minor relative to genuinely needed rubric content.

4 / 5

Actionability

Fully executable guidance: copy-paste-ready `case "$path" in ... esac` blocks for the exclusion gate and the path floor, an explicit combination rule (`severity = max(base_tier, path_floor, escalator_minimums)` over `low < medium < high < critical`), specific escalator shapes with exact minimums, and a byte-exact machine-readable output format (`## Severity: <critical|high|medium|low>`). It is not level 4 because no key execution detail is missing — the common cases are covered by the tier table plus escalator table plus the tricky-case golden set.

5 / 5

Workflow Clarity

The multi-step process is clearly sequenced (Step 1 axes with reachability cap → Step 2 with the exclusion gate explicitly run first, then floor, then escalators → Step 3 max-combination), with validation checkpoints in the output format (each escalator must cite "file:line", exclusion and floor matches shown as ✓/✗) and a feedback loop via the regression-tested golden eval suite that locks tricky cases (test-path destructive statement, billing path floor, dead-path catastrophic bug). The destructive/batch cap does not apply — the skill is advisory and performs no destructive or batch operations. Not level 4: checkpoints are explicit, not implicit.

5 / 5

Progressive Disclosure

Internal structure is good (a Contents TOC with anchors, and the rubric section explicitly flagged "self-contained"), but the body references files that do not exist in the skill bundle — `scripts/eval/golden/severity-tiering.jsonl`, `scripts/eval/l2.mjs`, `scripts/record-comment-relevance.mjs`, `agents/shared/rules/conventional-comments.md`, `rubric-composition.md`, `review-config.md` — and no references/, scripts/, or assets/ directories are present, so the references are unresolvable. Additionally, the consumer-integration material (the blocking-flag crosswalk and caller-consumption sections) is inlined where it could live in a reference file. This matches the level-3 anchor (some structure, references present but not clearly resolvable, content that should be separate is inline); not level 4 because dangling paths are more than a minor organization gap.

3 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete and comprehensive about what the skill does (tier emission, deterministic path floor, heuristic escalators, production-only scoping), explicit about when to use it, rich in natural trigger phrases with synonyms, and explicitly disambiguated from the adjacent confidence skill. Third-person voice throughout, dense without padding.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with comprehensive coverage: "Emits a lowercase tier (critical / high / medium / low) from a fixed axis rubric, then applies one deterministic, executable path floor (auth / billing / migration / infra / secrets paths) plus heuristic escalators (data-loss, security, concurrency shapes) — both scoped to reachable production code". It matches the anchor for multiple specific concrete actions; it is above level 4 because coverage of the skill's behavior is complete rather than having minor gaps.

5 / 5

Completeness

Both halves are explicit: what ("Rates how severe a finding or bug is if it is real — the blast-radius axis... Emits a lowercase tier... applies one deterministic, executable path floor... plus heuristic escalators") and when ("Use to triage a review finding or a bug. Triggers on..."), with concrete trigger phrases. This matches the anchor that clearly and explicitly answers both what AND when.

5 / 5

Trigger Term Quality

Trigger phrases are natural and varied: "how severe", "rate severity", "severity check", "triage this finding", "how bad is this", "/severity" — including synonyms ("how bad is this", "triage") and a slash-command form. Comprehensive coverage of natural phrasings; not level 4, which is for coverage missing a few natural terms.

5 / 5

Distinctiveness Conflict Risk

It carves a clear niche and explicitly disambiguates from the nearest neighbor: "complementary to confidence's is-it-real axis" and "the skill stays policy-free". Triggers are severity-specific, giving minimal conflict risk; it is not level 4 because the description actively distinguishes itself from closely related skills rather than merely having distinct triggers.

5 / 5

Total

20

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 3 missing, 2 deeper-than-1-level

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.