CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/hypothesis-testing

Authors property-based tests in Python using Hypothesis - wires `@given` with `strategies` (`st.integers`, `st.text`, `st.lists`, `st.from_regex`, `st.composite`), uses `assume()` / `.filter()` for preconditions, configures via `@settings(max_examples=..., deadline=...)`, and exploits Hypothesis's automatic shrinking to find the falsifying example. Integrates with pytest fixtures + parametrize. Use when a Python project needs PBT to catch edge cases the example-based tests miss - bug clusters around input ranges / boundary values / interaction between fields.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, mostly executable tutorial: clear step sequence, tight code examples, a genuine failure-to-regression feedback loop, and a clean one-level split with a real, comprehensive reference file. The residual costs are small: repeated attribution lines, occasional restating of the cited quickstart, and two placeholder/elided code spots.

DimensionReasoningScore

Conciseness

Mostly efficient: every step pairs a short rationale with tight, commented code, and it delegates the full strategies catalog, settings table, and anti-patterns to the reference file. Not 5 because of repeated "Per [hyp-quickstart]" attribution lines (4 occurrences) and a few explanations of things the cited doc already states, e.g. "The `draw()` call requests a value from a strategy; the composite returns the constructed value". Not 3 because the padding is minor and localized, not whole padded sections.

4 / 5

Actionability

Steps 1-8 give copy-paste-ready executable code (install command, property tests, composite strategy, filter/assume, settings, round-trip and metamorphic tests, pytest fixture), plus a concrete CI yaml snippet. Not 5 because Step 9's regression example uses a placeholder `@given(...)` rather than executable code, and the Step 8 fixture body is elided with "# ... setup ..." — minor gaps in an otherwise executable body.

4 / 5

Workflow Clarity

A clear, well-sequenced progression (install -> basic test -> strategies -> composites -> filtering -> settings -> property patterns -> pytest -> CI), with an error-recovery loop: on failure, copy the falsifying example into `@example` to lock in the regression. The "Heavy filtering is a smell - if 90% of generated cases are discarded, redesign the strategy" checkpoint guards the main fragility. Not 5 because there is no explicit validation step for CI integration (e.g., verifying derandomization actually reproduces the failure) and checkpoints are implied rather than enumerated. Not 3 because the sequence is coherent, gaps are minor, and no destructive/batch operations are involved.

4 / 5

Progressive Disclosure

The body is a lean tutorial that keeps only the getting-started subset inline (6 common strategies in Step 3, the 2 most important settings in Step 6) and points to a single real, one-level-deep reference (references/hypothesis-reference.md, verified to exist and contain the full catalog, settings table, anti-patterns, limitations) — signaled at Step 3, Step 6, and again in the References section. Cross-links to sibling skills are grouped for discovery. This matches the 5 anchor: clear overview, well-signaled one-level references, appropriate split. Not 4 because there is no content inlined that clearly belongs in the reference file.

5 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: comprehensive, concrete capability list in third-person voice with an explicit and specific 'Use when...' trigger clause. The only weakness is modest synonym coverage (fuzzing, generative testing) that slightly narrows the natural-language triggers.

DimensionReasoningScore

Specificity

Lists multiple concrete actions with the exact API surface: "wires `@given` with `strategies` (`st.integers`, `st.text`, ...)", "uses `assume()` / `.filter()` for preconditions", "configures via `@settings(max_examples=..., deadline=...)", "exploits... automatic shrinking", "Integrates with pytest fixtures + parametrize". Coverage spans authoring, filtering, configuration, shrinking, and test-runner integration — comprehensive. Not 4 because there are no meaningful gaps in capability coverage.

5 / 5

Completeness

Explicitly answers both: what ("Authors property-based tests in Python using Hypothesis... shrinking") and when ("Use when a Python project needs PBT to catch edge cases the example-based tests miss - bug clusters around input ranges / boundary values / interaction between fields"). The when-clause gives concrete trigger conditions, matching the 5 anchor. Not 4 because the trigger is fully specific, not merely present.

5 / 5

Trigger Term Quality

Includes natural phrases users would say: "property-based tests", "Python", "Hypothesis", "pytest", "edge cases", "boundary values", "PBT". Not 5 because it misses common synonyms such as "fuzzing"/"fuzz tests", "generative testing", or "invariant testing" that users often use interchangeably for this skill.

4 / 5

Distinctiveness Conflict Risk

"Python using Hypothesis" plus the library-specific API names carve a clear niche with minimal conflict risk — even sibling PBT skills (fast-check, jqwik, proptest, quickcheck) are for other languages and this description scopes to Python explicitly. Not 4 because the combination of language + library + testing paradigm leaves virtually no ambiguity about which skill applies.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents