CtrlK
BlogDocsLog inGet started
Tessl Logo

test-discipline

Design assertions that prove the actual contract and fail on realistic regressions

56

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.copilot/skills/test-discipline/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is lean, well-structured, and offers concrete actionable patterns without over-explaining. Its main weakness is that the guidance is a principles list rather than a sequenced workflow with validation checkpoints.

Suggestions

Add a short numbered workflow (e.g. identify the changed contract, write assertions naming the artifact, mutation-test the gate, verify the assertion turns red) to raise workflow clarity.

Include one minimal executable example (a before/after assertion or a mutation command) to push actionability toward 5.

Make the validation feedback loop explicit in the mutation-testing pattern ('change source, run test, confirm red, revert').

DimensionReasoningScore

Conciseness

Lean bullet-based body with no padding and no explanation of concepts Claude already knows; every section (Context, Patterns, Examples, Anti-Patterns) earns its place. Not below 5 because there is no verbosity or over-explanation to trim.

5 / 5

Actionability

Provides concrete actionable guidance ('Mutation-test critical gates by changing the real source', 'Make assertions name and inspect the offending artifact') with no pseudocode, but stops short of copy-paste-ready commands/examples. Not a 5 because there is no executable code or command; not a 3 because the guidance is specific and executable in intent rather than abstract.

4 / 5

Workflow Clarity

Content is a collection of principles/patterns rather than a sequenced multi-step workflow, with no explicit validation checkpoints. Not a 4 because there is no ordered sequence or feedback loop; not a 2 because the sections (Context, Patterns, Examples, Anti-Patterns) still give a coherent conceptual flow.

3 / 5

Progressive Disclosure

Under 50 lines with no external references needed and clearly organized into labeled sections, meeting the simple-skill exception for a top score. Not below 5 because the structure is clean and nothing belongs in a separate file.

5 / 5

Total

17

/

20

Passed

Description

50%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names concrete actions but omits any explicit 'when to use' trigger, which caps completeness. It is competent but would benefit from explicit trigger phrasing and broader keyword coverage.

Suggestions

Add an explicit 'Use when...' clause naming trigger situations (e.g. 'Use when changing an API, public interface, or workflow contract, or when strengthening regression tests').

Broaden trigger keywords with synonyms users actually say ('regression tests', 'test failures', 'mutation testing', 'contract tests').

Tighten the 'what' by pairing the two actions with the specific artifacts inspected, to push specificity toward 4-5.

DimensionReasoningScore

Specificity

Names the test-assertion domain and two concrete actions ('prove the actual contract', 'fail on realistic regressions') but coverage is not comprehensive. Not a 4 because it stops at two actions with no further specifics; not a 2 because the actions are concrete rather than generic.

3 / 5

Completeness

Gives a clear 'what' but no explicit 'when' / 'Use when' trigger guidance, so per the rubric guideline completeness is capped at 3. Not a 4 because the 'when' is entirely absent rather than weakly implied; not a 2 because the 'what' is clear and specific.

3 / 5

Trigger Term Quality

Includes relevant natural terms ('assertions', 'contract', 'regressions') but misses common variations/synonyms a user might say (e.g. 'regression tests', 'test failures'). Not a 4 because keyword coverage is thin and lacks 'Use when' phrasing; not a 2 because the terms present are natural rather than purely generic.

3 / 5

Distinctiveness Conflict Risk

Carves a test-discipline niche around contract-proving assertions but could still overlap with general testing/quality skills. Not a 4 because the trigger surface is narrow enough to risk confusion with broader testing skills; not a 2 because it is not broadly generic.

3 / 5

Total

12

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
bradygaster/squad
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.