CtrlK
BlogDocsLog inGet started
Tessl Logo

antithesis-research

Analyze a codebase to figure out how it should be tested with Antithesis: map the system, identify failure-prone areas and testable properties, and produce the research artifacts needed for workload and environment planning.

61

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./antithesis-research/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured orchestration skill: clear workflows with strong validation and feedback loops, actionable step-by-step guidance, and a clean reference table pointing at real one-level-deep files. The main costs are token weight from the inlined Self-Review checklist and duplicated output lists, plus one orphaned reference file.

Suggestions

Move the Self-Review checklist into a reference file (e.g. `references/self-review.md`) and keep a two-line pointer plus the most critical criteria in SKILL.md, cutting substantial context overhead.

Reference `references/scratchbook-artifacts.md` from the Reference Files table (or remove it) so every bundle file is reachable from the overview.

Deduplicate the artifact list between "Purpose and Goal" and "Output" — keep the success-criteria framing in one place and the flat output inventory in the other.

DimensionReasoningScore

Conciseness

The body is mostly tight and operational, but the ~30-line Self-Review checklist, the extended external-references threading explanation in Prerequisites, and the output list repeated across "Purpose and Goal" and "Output" are padding that could be tightened or moved to a reference. Fits anchor 3 (mostly efficient, some unnecessary explanation or could be tightened), below anchor 4 because the duplication and inlined checklist are more than minor trims.

3 / 5

Actionability

Guidance is largely executable for an instruction-only skill: exact artifact paths, precise SDK assertion names to scan for (`assert_always!`, `assert_sometimes!`, …) with what to record per hit, and concrete downstream-consumer expectations. Score 4 rather than 5 because the assertion scan (step 4) describes what to search for but gives no copy-paste grep/ripgrep command, and some steps defer entirely to reference files.

4 / 5

Workflow Clarity

Three clearly sequenced workflows with numbered steps, an explicit evaluation feedback loop ("apply refinements, fill gaps, escalate biases to the user"), and a closing Self-Review checklist including a fresh-context reviewer for blind spots. This matches anchor 5 (explicit validation steps, feedback loops, checklists); not 4 because validation is not merely present but multi-layered.

5 / 5

Progressive Disclosure

SKILL.md works as an overview: a Reference Files table with a "When to read" column signals one-level-deep references clearly, and every referenced path resolves to a real file in references/. Two minor gaps hold it at anchor 4 rather than 5: `references/scratchbook-artifacts.md` exists in the bundle but is referenced nowhere (orphaned navigation-wise), and the long Self-Review section is content that plausibly belongs in a reference file.

4 / 5

Total

16

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description with concrete actions and a clearly distinct Antithesis niche. Its one real weakness is the absence of any "Use when…" trigger clause, which caps completeness and weakens its value for skill selection.

Suggestions

Append an explicit trigger clause, e.g. "Use when starting Antithesis testing work on a codebase, when planning workloads or test properties, or when the user mentions Antithesis, testing research, or finding invariants."

Name the concrete outputs in the description (e.g. "a system analysis, property catalog, and deployment topology") instead of the generic "research artifacts".

Add one or two natural user phrasings as trigger synonyms, such as "find invariants" or "what should we test", to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

Lists several specific actions — "map the system", "identify failure-prone areas and testable properties", "produce the research artifacts needed for workload and environment planning" — but "research artifacts" stays generic rather than naming the concrete outputs. This sits between anchor 3 (1-2 concrete actions) and anchor 5 (comprehensive, fully concrete action list).

4 / 5

Completeness

The "what" is clearly and concretely answered, but there is no "Use when…" clause or equivalent explicit trigger guidance — "needed for workload and environment planning" hints at downstream context, not when to invoke the skill. Per the judging guidelines, a missing 'Use when' clause caps completeness at 3.

3 / 5

Trigger Term Quality

Good natural keyword coverage — "codebase", "tested", "failure-prone areas", "testable properties", plus the distinctive "Antithesis" — but misses common user phrasings like "testing research", "property-based testing", or "find invariants". Fits anchor 4 (good coverage, a few natural terms missing), not 5 (no synonym/variation set) and not 3 (keywords are more than merely 'some relevant').

4 / 5

Distinctiveness Conflict Risk

Clear niche (Antithesis-specific test research) with distinctive triggers like "failure-prone areas and testable properties" that no general testing or analysis skill would claim. Matches anchor 5; it is not merely 'mostly distinct' (anchor 4) because the tool name alone disambiguates it.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
antithesishq/antithesis-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.