CtrlK
BlogDocsLog inGet started
Tessl Logo

antithesis-review-inputs

Review a workload's input generation for state space exploration: how well it uses randomness to drive the system under test into diverse regions of behavior.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./antithesis-review-inputs/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable review procedure with excellent progressive disclosure — the single reference file is real, one level deep, and clearly motivated. The main residual improvements are tightening the repeated tone framing and operationalizing the user-correction step in the report workflow.

DimensionReasoningScore

Conciseness

The body is efficient and domain-specific — no padding explaining concepts Claude already knows — but the framing 'a workload exhibiting anti-patterns may still be productive/finding bugs' is restated in both the Purpose and Tone sections, and Tone could merge into fewer words. This is 'efficient; minor instances of over-explanation that could be trimmed'.

4 / 5

Actionability

Fully concrete guidance for an instruction-only skill: a six-step workflow, specific examination checklists ('Are there symmetric pairs always active together?', 'Do payload sizes span orders of magnitude?'), tier definitions with examples, and an explicit good-vs-bad example of a finding ('Your three actions (put, get, delete) are always all active with equal weight...' beats 'all actions are always active').

5 / 5

Workflow Clarity

Six clearly numbered, well-sequenced steps with an ask-the-user checkpoint ('If the workload location isn't obvious, ask the user') and a present-your-model-first feedback mechanism. It falls short of a 5 because the model-correction step ('presented so the user can correct misunderstandings before reading findings') is described rather than operationalized as an explicit validation checkpoint, and there are no other verification steps.

4 / 5

Progressive Disclosure

SKILL.md is a genuine overview and all anti-pattern detail is correctly delegated to a single one-level-deep reference, clearly signaled in workflow step 2 ('Read references/anti-patterns.md before examining the workload') and reinforced by a 'when to read' table. The reference file exists and matches its description (a full catalog organized by Action Selection, Data Generation, and Structural/Timing categories).

5 / 5

Total

18

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is precise about what the skill does and its domain, but it omits any when-to-use guidance and lacks the natural trigger phrases that would help a user (or the harness) select it. Adding a 'Use when...' clause with user-facing phrasing would address both gaps at once.

Suggestions

Add an explicit 'Use when...' clause to the description, e.g. 'Use when a user isn't finding bugs and wants to understand whether the workload's input generation is limiting what Antithesis can explore.'

Include natural trigger terms users would actually say — 'not finding bugs', 'health check', 'workload review' — alongside the existing domain keywords.

Mention 'Antithesis' explicitly in the description to sharpen distinctiveness and reduce conflict risk with generic test-review skills.

DimensionReasoningScore

Specificity

The description names the domain concretely ("Review a workload's input generation for state space exploration") and the mechanism ("how well it uses randomness to drive the system under test into diverse regions of behavior"), but it lists only that single composite review action rather than several specific actions, matching the 'names domain and 1-2 concrete actions, but not comprehensive' anchor.

3 / 5

Completeness

The 'what' is clear and concrete, but there is no 'Use when...' clause or equivalent trigger guidance anywhere in the description, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Relevant domain keywords are present ("workload", "input generation", "randomness", "state space"), but common natural phrasings a user would actually say — "not finding bugs", "health check", "fuzzing" — are absent from the description despite appearing in the body's When-to-Use section. This is 'some relevant keywords but missing common variations or synonyms'.

3 / 5

Distinctiveness Conflict Risk

A mostly distinct niche (reviewing randomness-driven input generation for state space coverage) with only minor overlap risk against general test-review or code-review skills; curiously, the description never mentions 'Antithesis' itself, which would sharpen the niche further.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
antithesishq/antithesis-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.