CtrlK
BlogDocsLog inGet started
Tessl Logo

antithesis-triage

Triage Antithesis test reports to understand what happened in a run: look up runs, check status, investigate failed properties (assertions), view metadata, download logs, inspect findings, and examine environmental details. Load after a run completes or when investigating a failure.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured triage skill body: validated preflight, clearly sequenced workflows with edge-case handling, disciplined progressive disclosure, and a self-review checklist. The remaining slack is minor — some duplicated framing text and a few steps that defer detail to references instead of naming the exact command inline.

Suggestions

Drop the redundant "Use this skill to analyze Antithesis test runs." line and the duplicate "Skill version" line (the frontmatter already carries both) to tighten the body.

Compress the verbose property-investigation step e (comparing counterexample logs) into a short bullet list of the comparison heuristics.

Inline the one-line form of the log-download command next to the reference pointer in the triage workflow so the most common path is executable without a file read.

DimensionReasoningScore

Conciseness

The body is efficient — commands, conditions, and workflows with little concept re-teaching — but has minor trimmable fat: "Use this skill to analyze Antithesis test runs." duplicates the description, the version string appears twice (frontmatter and body), and the property-investigation step e paragraph is wordy. Fits the 4 anchor ('minor instances of over-explanation'), not 5.

4 / 5

Actionability

Copy-paste-ready commands are present ("snouty doctor --json", the jq tenant pipeline, "snouty runs --json events ${RUN_ID} ${PROPERTY_NAME}", "snouty runs --json build-logs ${RUN_ID}") with explicit stop conditions, but run discovery, log download, and property parsing are deferred to reference files rather than shown — mostly executable with minor gaps, the 4 anchor rather than fully covering common cases in-body.

4 / 5

Workflow Clarity

Workflows are clearly sequenced with explicit validation and error recovery: the preflight gate ("Proceed only when the top-level `ok` is `true` and the `api_key` check's `status` is `ok`. Otherwise relay the failing check's `message`/`notes` and stop"), 404 handling for non-triageable runs, the incomplete-run fallback (failure_moment + build-logs), and a closing Self-Review checklist — matching the 5 anchor (explicit validation steps, feedback loops, checklists).

5 / 5

Progressive Disclosure

The body states an explicit disclosure policy ("Do NOT read them all up front — only read a reference file when you are told to. Each reference file is mentioned by name at the point where it is needed") and all five references/ files cited by name (run-discovery.md, run-info.md, properties.md, logs.md, instrumentation.md) exist and are invoked at their point of need, one level deep — the 5 anchor.

5 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, comprehensive action list in third person with an explicit load/trigger clause. The only weakness is mild overlap risk with the sibling Antithesis debugging skills around the phrase "investigating a failure".

DimensionReasoningScore

Specificity

The description lists seven concrete actions ("look up runs, check status, investigate failed properties (assertions), view metadata, download logs, inspect findings, and examine environmental details"), giving comprehensive coverage of the triage domain with no notable gaps, matching the top anchor.

5 / 5

Completeness

It explicitly answers what (the enumerated triage actions) and when ("Load after a run completes or when investigating a failure"), matching the anchor requiring both with concrete trigger phrases; not 4, since the when-clause is explicit rather than improvable.

5 / 5

Trigger Term Quality

Natural user phrasing is well covered — "test reports", "check status", "download logs", "investigating a failure" — and includes synonyms such as "properties (assertions)" and "findings", fitting the comprehensive-synonyms anchor rather than the 'a few natural terms missing' anchor at 4.

5 / 5

Distinctiveness Conflict Risk

"Triage Antithesis test reports" carves out a clear niche, but "investigating a failure" would also naturally route to the closely related antithesis-debug skill, so there is minor overlap risk with sibling skills — the 4 anchor, not the minimal-conflict 5 anchor.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
antithesishq/antithesis-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.