CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/spec-testability-heuristics

Judges whether a written requirement can be tested at all, by scoring each claim against three heuristics: Observable (a state or output visible from outside the system), Decidable (a deterministic pass/fail against a named oracle), and Bounded (stated inputs, states, and users). Supplies untestable-to-testable rewrite pairs per heuristic, severity assignment, a BLOCK / REVIEW / OK verdict rule, a findings table that pairs every flagged sentence with a concrete replacement, and the review workflow for running the rubric against a story, PRD section, PR description, or spec at sprint planning or PR review, with per-verdict hand-off targets. Scope is testability only: it does not write the tests, rank risk, or judge whether the requirement is correct or complete. Use when a user story, acceptance criterion, PRD section, API contract, or PR description is about to be handed to implementation and someone needs to know which sentences cannot be verified as written.

77

Quality

97%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality, highly actionable SKILL.md body with a clear workflow, explicit validation checkpoints, and appropriate offloading of worked examples. The only weakness is moderate verbosity from ISTQB glossary citations and concept primers that restate knowledge Claude already has.

Suggestions

Drop or compress the inline ISTQB glossary quotations (testability, expected result, test oracle, equivalence partitioning, boundary value analysis); cite the glossary once and rely on Claude's existing knowledge of these terms to save several hundred tokens.

Trim the 'What testability means here' shift-left rationale to one or two sentences — the cost-benefit framing restates a concept Claude already understands and is not needed to run the rubric.

Keep the operational sections (heuristics, rewrite tables, severity, verdict, output format) exactly as-is; they are the lean, high-signal core and should be preserved.

DimensionReasoningScore

Conciseness

The operational core (heuristics, failure signatures, rewrite tables, severity, verdict, workflow, output format, anti-patterns, limitations) is lean and high-signal, but recurring ISTQB glossary citations and the shift-left / equivalence-partitioning primers explain concepts Claude already knows and could be trimmed.

4 / 5

Actionability

Provides fully executable guidance: per-heuristic failure signatures with trigger words, untestable-to-testable rewrite tables, explicit rewrite-move formulas, a six-step procedure, a severity table, a verdict rule, and a copy-paste-ready output format template with an example findings table.

5 / 5

Workflow Clarity

Clearly sequenced 'How to apply the rubric' and 'Review workflow' with explicit checkpoints — 'record every heuristic it fails', 'a finding with no proposed replacement is not a finding', verdict-from-severities — plus a hand-off-target map and an anti-patterns checklist.

5 / 5

Progressive Disclosure

The overview keeps the operational heuristics and tables inline where a reviewer needs them, and offloads only the worked examples to a well-signaled, one-level-deep reference (references/worked-examples.md) that exists in the bundle.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person voice, comprehensive concrete capabilities, explicit 'Use when' trigger with natural artifact terms, and a sharp scope boundary that distinguishes it from related QA skills. No substantive weaknesses.

DimensionReasoningScore

Specificity

Names the domain (requirement testability) and lists multiple concrete actions — scoring claims against three named heuristics, rewrite pairs, severity assignment, a verdict rule, a findings table, a review workflow, and per-verdict hand-off targets — giving comprehensive coverage.

5 / 5

Completeness

Explicitly answers both 'what' (judges testability via three heuristics and supplies rewrites, verdict rule, findings table, workflow) and 'when' via a concrete 'Use when a user story, acceptance criterion, PRD section, API contract, or PR description is about to be handed to implementation' trigger clause.

5 / 5

Trigger Term Quality

Includes natural terms a user or PM would actually say — 'user story', 'acceptance criterion', 'PRD section', 'API contract', 'PR description', 'spec', 'sprint planning', 'PR review' — with good variation across artifact types.

5 / 5

Distinctiveness Conflict Risk

Carves a clear niche (testability-only review) and explicitly fences off adjacent jobs ('it does not write the tests, rank risk, or judge whether the requirement is correct or complete'), minimizing overlap with sibling skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents