Review a workload's input generation for state space exploration: how well it uses randomness to drive the system under test into diverse regions of behavior.
60
70%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./antithesis-review-inputs/SKILL.mdSkill version: 2026-09-23 a8fe900
Examine a workload and report how well its input generation explores state space. The system under test is a black box — the review works entirely from the workload code and any documentation the user provides.
This skill is read-only. It produces a report of findings. It does not modify the workload, open PRs, or suggest specific code changes.
The framing is "untapped exploration" — findings are ceilings on what the workload can reach, not defects. A workload exhibiting any anti-pattern may still be productive. The question is whether it could reach more interesting behavior with structural changes to how it uses randomness.
When a user isn't finding bugs and wants to understand whether the workload's input generation is limiting what Antithesis can explore. Also useful as a periodic health check on workloads — both human-written and agent-generated.
Randomize the policy per run, not just the choices under a fixed policy. Each run should explore a different region of behavior by biasing its configuration differently. The skill reviews how well the workload does this.
Find the test commands and workload code. Common locations:
antithesis/test/ — test command scriptsIf the workload location isn't obvious, ask the user.
Read references/anti-patterns.md before examining the workload. It contains
the full catalog of input generation anti-patterns organized by category, with
explanations of why each limits exploration and what better looks like.
Read the workload code and look for structural patterns that match the anti-patterns. Focus on what is visible in the code:
Action selection — How are operations chosen?
Data generation — How are inputs produced?
Structure — How is the workload organized?
For each anti-pattern you identify in the workload:
The explanation should be concrete and specific to this workload, not a generic restatement of the anti-pattern. "Your three actions (put, get, delete) are always all active with equal weight — runs where only deletes fire would test behavior with zero keys" is better than "all actions are always active."
Not all findings deserve equal weight. Tier them:
When you can't assess a finding's severity without domain knowledge (e.g., "this range is hardcoded, but I don't know what the system accepts"), say so. Present the structural fact and let the user supply the domain judgment.
Structure the report as:
What the workload does — A brief summary of what the reviewer sees: what actions the workload takes, how it generates data, how it's structured. This is the reviewer's model of the workload, presented so the user can correct misunderstandings before reading findings.
Findings — Organized by tier (high signal first), grouped by category when multiple findings cluster. Each finding names the pattern, points to the code, and explains the unexplored state space in terms specific to this workload.
What's working well — Patterns the workload already does right. If the workload has per-run action subsetting, or varies payload sizes, say so. Findings-only reports feel like a critique; acknowledging what works frames the report as an assessment.
The report must read as "here's what you could explore that you're not reaching" — never as "your workload is bad." A workload exhibiting multiple anti-patterns may still be finding bugs; the goal is to help it find more.
Don't frame findings as "you need to do this." Frame them as "this is state space you're not reaching, and here's what reaching it would look like." The user decides what's worth changing.
This skill reviews input generation — how the workload drives the system into diverse states. It does not review:
If the reviewer notices something outside its scope that seems important (e.g., the workload has no assertions at all), it can mention it briefly as an observation, but should not review it in depth.
| Reference | When to read |
|---|---|
references/anti-patterns.md | Before examining any workload — the full catalog |
1fd8470
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.