CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/quality-status-digest

Computes a recurring quality status digest from metrics that already exist: CI pass rate with an explicit denominator rule, escape-defect count, and a flake-debt score, assigns red / amber / green per area against stated thresholds, then rolls the same per-team rows into a portfolio view with a severity-by-blast-radius heatmap, STABLE / WATCH / INVEST tags, and a capacity flag. Keeps DORA delivery metrics separate from defect-leakage and flake measures instead of blending them under one label. Produces the status artifact only: it does not instrument anything, does not define SLOs or targets, and does not decide what gets fixed first. Use when a weekly quality review, sprint check-in, or quarterly portfolio review is due and the CI history, defect tracker, and quarantine list already hold the numbers but nobody has assembled them into one page.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered two-part instruction skill: every formula, cut point, and tagging rule is stated with its inputs, its fallback when data is missing, and an honest label of which numbers are conventions. The only weakness is mild redundancy between the frontmatter description and the body's Overview/ownership sections, plus the DORA-separation point repeated across four locations.

Suggestions

Trim the 'What this skill owns, and what it does not' section to a short bullet list or fold it into the Overview, since the frontmatter description already states the non-goals (no instrumentation, no SLOs, no prioritization).

State the DORA-separation rule once (Step 2, where it is introduced and linked to references/dora-metrics.md) and reference it from Step 6 and Limitations instead of re-arguing it in each place.

DimensionReasoningScore

Conciseness

The body is dense with domain judgment Claude would not already have — denominator rules, threshold rationales ('Below ~90%, a failing build stops being a signal and engineers re-run by reflex'), convention caveats ('practitioner conventions, not standards'), and an anti-patterns table — so most tokens earn their place. Minor trim opportunities exist: the Overview and the 'What this skill owns, and what it does not' section restate the frontmatter description's scope and non-goals at length, and the DORA-separation point is made in Step 2, Step 6, Limitations, and the reference file.

4 / 5

Actionability

Guidance is fully executable for an instruction-only skill: exact formulas ('pass_rate = successful_runs / (successful_runs + failed_runs)', 'flake_debt_score = (stale_quarantine_count * 2) + new_flakes_in_window'), a complete numeric threshold table, concrete output paths ('quality-digest/<YYYY-MM-DD>.md'), a specified machine-readable 'digest-row' line, and full worked templates in references/output-templates.md that exist and match what the body promises. Missing-data fallbacks are explicit (report raw count with unknown denominator; '[DATA NOT SUPPLIED]' cells).

5 / 5

Workflow Clarity

Ten steps are sequenced across two clearly separated altitudes (Part 1 Steps 1–5, Part 2 Steps 6–10) with an explicit cross-part contract ('the final fenced digest-row line is what the portfolio roll-up consumes'). Edge cases get explicit handling at each point they arise — unavailable deployment denominator, blast-radius data not supplied, first portfolio review with no prior period ('Say so instead of tagging everything STABLE by default') — and the anti-patterns table with fix references acts as an error-prevention checklist.

5 / 5

Progressive Disclosure

SKILL.md is a navigable overview that keeps verbatim definitions and full templates out of the main file in two real, one-level-deep references (references/dora-metrics.md, references/output-templates.md), both existing and both clearly signaled from the exact steps that need them (Steps 2, 5, 6, 10). No nested references; bulk detail is appropriately split and easy to find.

5 / 5

Total

19

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description in third-person voice: comprehensive concrete capabilities, an explicit and scenario-rich 'Use when' clause, and non-goals that fence off adjacent skills. The only gap is a handful of natural synonym phrases a user might use instead of the stated ones.

DimensionReasoningScore

Specificity

Multiple specific concrete actions with comprehensive coverage: 'Computes a recurring quality status digest… CI pass rate with an explicit denominator rule, escape-defect count, and a flake-debt score, assigns red / amber / green per area against stated thresholds, then rolls the same per-team rows into a portfolio view with a severity-by-blast-radius heatmap, STABLE / WATCH / INVEST tags, and a capacity flag.' Nothing is abstract; each capability is named.

5 / 5

Completeness

Explicitly answers both questions: the 'what' is the full computation-and-roll-up sentence above, and the 'when' is a concrete trigger clause — 'Use when a weekly quality review, sprint check-in, or quarterly portfolio review is due and the CI history, defect tracker, and quarantine list already hold the numbers but nobody has assembled them into one page.'

5 / 5

Trigger Term Quality

Good natural-phrase coverage: 'weekly quality review', 'sprint check-in', 'quarterly portfolio review', 'CI history', 'defect tracker', 'quarantine list' are phrases a real user would say. A few natural variations are missing (e.g. 'QA report', 'test quality summary', 'flaky tests' as a plain-language synonym for flake debt), so it sits just below the comprehensive-with-synonyms anchor rather than at it.

4 / 5

Distinctiveness Conflict Risk

A clear niche (assembling existing metrics into a recurring status artifact) plus explicit boundary statements — 'it does not instrument anything, does not define SLOs or targets, and does not decide what gets fixed first' and 'Keeps DORA delivery metrics separate from defect-leakage and flake measures' — that sharply reduce the odds of triggering for an instrumentation, SLO-setting, or prioritization skill.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents