CtrlK
BlogDocsLog inGet started
Tessl Logo

antithesis-debug

Interactively debug an Antithesis test run in the multiverse debugger (MVD): launch a session from a run, open a debugging-session URL, and inspect container filesystem and runtime state from inside the run.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary body: executable commands, validated multi-step workflows with error-recovery guidance, and a clean one-level reference structure over real bundle files. The only meaningful cost is token efficiency — the runtime-reinjection and page-loading guidance is repeated across sections and could be consolidated into one place.

Suggestions

Consolidate the runtime-reinjection guidance, which appears in "Runtime injection", "General guidance", and "Advanced mode only" — state it once and reference that spot, saving ~15-20 lines.

Trim or fold the "Page loading checks" section, since waitForReady usage and its return shape largely duplicate what the workflows and reference files already cover.

The one-shot probes and diagnostics code blocks (loadingFinished/loadingStatus in both namespaces) could move to a reference file, keeping the SKILL.md overview leaner.

DimensionReasoningScore

Conciseness

Nearly all content is Antithesis-specific (mode detection via getMode(), the CAMPAIGN SAW TERMINAL EVENT failure mode, the tab-strip vs switchMode() disagreement) that Claude could not know otherwise, with no generic-concept padding. However, runtime-reinjection guidance repeats three times ("Runtime injection", "General guidance", "Advanced mode only") and "Page loading checks" duplicates waitForReady info already in the workflows, fitting the 4 anchor's "minor instances that could be trimmed" rather than 5.

4 / 5

Actionability

Fully copy-paste-ready commands throughout: `cat assets/antithesis-debug.js | agent-browser --session "$SESSION" eval --stdin`, exact eval calls like `window.__antithesisDebug.simplified.runCommand("ls -la /")`, switchMode calls, and documented return shapes (`{ ok, ready, attempts, waitedMs }`) covering the common cases. Not below 5 — no pseudocode or vague steps anywhere.

5 / 5

Workflow Clarity

Four numbered workflows each with explicit "read X first" steps, waitForReady() validation checkpoints after navigation, explicit error-recovery loops (reinject the runtime when window.__antithesisDebug is missing; fork a fresh branch after a nonzero exit code), and a closing self-review checklist. No destructive or batch operations trigger the validation cap.

5 / 5

Progressive Disclosure

The body is an overview with when-to-read tables mapping each task to its reference file; all 7 referenced files in references/ and all 3 cited assets exist and are one level deep, with cross-links between sibling references rather than nested chains. Matches the 5 anchor: clear overview, well-signaled one-level references, easy navigation.

5 / 5

Total

19

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, domain-distinct, and written in third person with concrete actions, but it lacks any explicit "when to use" trigger clause, capping completeness. Adding a sentence of trigger conditions (e.g., when the user supplies a debugging-session URL or asks to inspect container state) would lift it to the top anchors.

Suggestions

Append a "Use when..." clause to the description, e.g. "Use when the user provides an Antithesis debugging-session URL or asks to inspect container filesystem, runtime state, or events inside a run" — this removes the completeness cap at 3.

Mention the two debugger modes (simplified shell commands vs. advanced JavaScript notebook) in the description so users searching for "notebook" or "run a command in a container" also match.

Consider mentioning event-log download/analysis, since it is a supported workflow in the body but absent from the description.

DimensionReasoningScore

Specificity

Names several concrete actions — "launch a session from a run", "open a debugging-session URL", "inspect container filesystem and runtime state" — but omits coverage the body delivers (simplified/advanced modes, event-log download, notebook workflows). It exceeds the 3 anchor (1-2 concrete actions) yet falls short of the 5 anchor's comprehensive coverage.

4 / 5

Completeness

The "what" is clear (launch, open URL, inspect state), but there is no "Use when..." clause or equivalent explicit trigger guidance in the description — the rubric's judging guideline caps completeness at 3 in that case. The trigger guidance exists only in the body's "When to use this skill" section, not the description.

3 / 5

Trigger Term Quality

Good natural term coverage for the domain — "Antithesis", "debug", "multiverse debugger", "MVD", "debugging-session URL", "container", "filesystem", "runtime state" — including the MVD acronym users would say. A few natural terms are missing ("events log", "notebook", "time travel"), so it fits the 4 anchor rather than 5.

4 / 5

Distinctiveness Conflict Risk

"Antithesis test run", "multiverse debugger (MVD)", and "debugging-session URL" define a clear niche with distinct triggers and minimal conflict risk with other skills. No neighboring anchor fits better.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
antithesishq/antithesis-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.