CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/spec-testability-heuristics

Judges whether a written requirement can be tested at all, by scoring each claim against three heuristics: Observable (a state or output visible from outside the system), Decidable (a deterministic pass/fail against a named oracle), and Bounded (stated inputs, states, and users). Supplies untestable-to-testable rewrite pairs per heuristic, severity assignment, a BLOCK / REVIEW / OK verdict rule, a findings table that pairs every flagged sentence with a concrete replacement, and the review workflow for running the rubric against a story, PRD section, PR description, or spec at sprint planning or PR review, with per-verdict hand-off targets. Scope is testability only: it does not write the tests, rank risk, or judge whether the requirement is correct or complete. Use when a user story, acceptance criterion, PRD section, API contract, or PR description is about to be handed to implementation and someone needs to know which sentences cannot be verified as written.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured rubric skill with excellent actionability: concrete rewrite pairs, exact failure vocabularies, a deterministic severity-to-verdict rule, and a complete output format. The main cost is token weight from repeated ISTQB glossary citations and motivational framing that Claude does not need; the workflow is clearly sequenced but lacks an explicit final verification checkpoint.

Suggestions

Trim the ISTQB glossary quotations (testability, expected result x2, test oracle, equivalence partitioning, boundary value analysis) to a single one-line citation or drop them entirely — Claude already knows these concepts, and the heuristics' own failure signatures carry the load.

Cut the 'shift left' paragraph and aphorisms ('The cheapest defect to fix is the one prevented before it's coded') to one sentence of placement guidance; the workflow section already says when to run the review.

Add an explicit final verification step to the workflow, e.g. 'Before emitting, confirm every non-OK row has a rewrite and the verdict matches the severity table', to close the last workflow-clarity gap.

DimensionReasoningScore

Conciseness

The body is mostly tight (failure-signature word lists, compact rewrite tables), but it repeatedly explains concepts Claude already knows: quoted ISTQB definitions of 'testability', 'expected result' (cited twice), 'test oracle', 'equivalence partitioning', and 'boundary value analysis', plus a 'shift left' rationale and aphorisms like 'The cheapest defect to fix is the one prevented before it's coded'. This matches 'mostly efficient but includes some unnecessary explanation or could be tightened'; it is not a 4 because the glossary quoting is a recurring pattern rather than a minor instance.

3 / 5

Actionability

For an instruction-only skill the guidance is fully concrete: three untestable-to-testable rewrite tables ('p95 page-load time on `/dashboard` is at most 2.5s under 100 concurrent users'), exact failure-signature vocabulary ('fast, clean, robust, seamless...'), a deterministic severity table, a verdict rule, and a filled-in output-format findings table. Not a 4: there are no gaps — every step produces a specified artifact with copy-paste-ready structure.

5 / 5

Workflow Clarity

A clear 6-step sequence with built-in checkpoints ('Record every heuristic it fails, not just the first', 'A finding with no proposed replacement is not a finding') and an anti-patterns table that serves as error-recovery guidance. Not a 5: there is no explicit final verification step (e.g., confirming the emitted verdict matches the severity table or that every flagged claim carries a rewrite before emitting), only implicit rules; not a 3 because the sequence and checkpoints are otherwise explicit and the operation is read-only, so the destructive/batch cap does not apply.

4 / 5

Progressive Disclosure

The single bundle file, references/worked-examples.md, exists, is linked with a descriptive signpost ('Three worked reviews - one BLOCK verdict..., one OK verdict..., and one REVIEW verdict... are in [references/worked-examples.md]'), and is exactly one level deep (it links back to SKILL.md, no further nesting). The rubric itself belongs in SKILL.md and the examples are appropriately split out. Not a 4: navigation and placement leave no real gaps.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person, information-dense, with concrete capability enumeration, an explicit 'Use when' trigger clause covering all natural artifact synonyms, and stated scope exclusions that sharply de-conflict it from adjacent test-writing and risk-ranking skills. No weaknesses found within the rubric's dimensions.

DimensionReasoningScore

Specificity

It lists multiple concrete capabilities with no padding-by-abstraction: 'scoring each claim against three heuristics: Observable... Decidable... Bounded', 'Supplies untestable-to-testable rewrite pairs per heuristic, severity assignment, a BLOCK / REVIEW / OK verdict rule, a findings table', and 'the review workflow... with per-verdict hand-off targets'. Matches 'lists multiple specific concrete actions; comprehensive coverage' and is well above anchor 4 since every clause names a distinct deliverable.

5 / 5

Completeness

Both halves are explicit: the 'what' opens the description ('Judges whether a written requirement can be tested at all, by scoring each claim against three heuristics') and the 'when' is an explicit 'Use when a user story, acceptance criterion, PRD section, API contract, or PR description is about to be handed to implementation'. It even adds scope exclusions ('it does not write the tests, rank risk, or judge whether the requirement is correct or complete'), exceeding anchor 5's requirement.

5 / 5

Trigger Term Quality

Natural user phrasing is comprehensively covered: 'user story, acceptance criterion, PRD section, API contract, or PR description', 'at sprint planning or PR review', and 'cannot be verified as written'. It includes the artifact synonyms users would actually name when handing off work; not a 4 because no commonly said variant of the trigger situation is missing.

5 / 5

Distinctiveness Conflict Risk

A clear niche with explicit de-confliction: 'Scope is testability only: it does not write the tests, rank risk, or judge whether the requirement is correct or complete', and named hand-off targets to sibling skills (gherkin-from-stories, non-functional-requirement-extractor) keep it from absorbing neighbors' triggers. Minimal conflict risk; not a 4 because the boundary is stated, not merely implied.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents