CtrlK
BlogDocsLog inGet started
Tessl Logo

test-audit

Invoke whenever writing, changing, reviewing, or sweeping tests in the eve repository. Authoring gate for new tests plus audit workflow for low-value, slow, implementation-coupled, or duplicative tests and the test-only production seams they demand.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally disciplined instruction-only skill: dense, executable, and sequenced with real validation gates for destructive batch operations, with no filler. The only structural weakness is the dangling CAMPAIGN.md reference given no reference files ship with the skill bundle.

Suggestions

Ship CAMPAIGN.md alongside SKILL.md (or inline a brief campaign-mode summary) so the body's only in-skill reference resolves.

Provide a concrete fallback for the `vitest.<tier>.config.ts` and `/tmp/<tier>.json` placeholders (e.g., naming the available tier configs) so the discovery command is runnable without guessing.

DimensionReasoningScore

Conciseness

The body is lean and telegraphic ("Three modes, one value bar", "Prefer net-negative production LOC") with zero re-teaching of concepts Claude already knows and no padded sections — every line adds project-specific rules, matching anchor 5 exactly. It is not a 4 because no section could be trimmed without losing a real constraint.

5 / 5

Actionability

Fully executable, copy-paste-ready commands appear throughout (the vitest JSON-reporter baseline command, `pnpm --filter eve exec vitest run --config vitest.<tier>.config.ts <path>`, `pnpm guard:invariants`, `pnpm fmt/lint/typecheck`), plus concrete checklists (four gate questions, required evidence fields, per-pattern junk list). This matches anchor 5; anchor 4 would require minor gaps in the commands or checklists, and none are evident.

5 / 5

Workflow Clarity

The workflow is clearly sequenced with explicit validation checkpoints and feedback loops: read-only discovery first, 'a missing field means the candidate is not ready for deletion', a numbered validation list with exact commands, and the rule that a failing retained test is treated as a product bug and reproduced rather than deleted. This is a destructive/batch skill whose validation is thorough, so the cap at 3 does not apply; it matches anchor 5's checklists and error-recovery loop.

5 / 5

Progressive Disclosure

Structure is good: a ~200-line overview that defers campaign detail to a clearly signaled one-level reference ("before starting one, read [CAMPAIGN.md](CAMPAIGN.md)") and points to AGENTS.md and the gh-pr-description skill. It is not 5 because the skill ships no bundle files and the referenced CAMPAIGN.md does not exist alongside SKILL.md, so that reference is dangling as delivered — a minor navigation gap rather than anchor 3's unclear or buried references.

4 / 5

Total

19

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, trigger-explicit description: the 'when' clause is exemplary and the domain is sharply scoped to the eve repository. The 'what' is specific but compressed and incomplete — campaign mode is absent from the description despite being one of the skill's three modes.

Suggestions

Mention campaign mode (pruning a subsystem's test surface) in the description so the 'what' covers all three modes the body defines.

Add natural trigger synonyms users are likely to say, such as 'delete', 'prune', or 'clean up tests', and 'flaky tests', to broaden trigger-term coverage.

Loosen the dense noun phrase ('low-value, slow, implementation-coupled, or duplicative tests and the test-only production seams they demand') into plainer verb-led statements of what the skill does.

DimensionReasoningScore

Specificity

The description names several concrete actions — "Authoring gate for new tests plus audit workflow for low-value, slow, implementation-coupled, or duplicative tests and the test-only production seams" — going well beyond the 1-2 actions of anchor 3, but it omits the third mode (campaign pruning) and relies on compressed noun phrases rather than enumerated capabilities, so it falls short of anchor 5's comprehensive coverage.

4 / 5

Completeness

Both are explicit: "Invoke whenever writing, changing, reviewing, or sweeping tests in the eve repository" answers when with concrete trigger phrases, and the authoring-gate/audit-workflow clause answers what. It is not a 5 because the 'what' is a dense noun stack that leaves campaign mode unstated, so the what/when pairing is clear but not maximally explicit.

4 / 5

Trigger Term Quality

Natural phrases users would say are present ("writing, changing, reviewing, or sweeping tests", "low-value", "slow", "duplicative tests"), giving good keyword coverage; common variations like "delete/prune/clean up tests", "flaky tests", or "test coverage" are missing, matching anchor 4 rather than anchor 5's synonym-complete coverage.

4 / 5

Distinctiveness Conflict Risk

Scoping to "the eve repository" and distinctive vocabulary ("test-only production seams", tiered test audit) give it a clear niche with minor overlap risk; it stops short of anchor 5 because it claims every act of writing or changing tests, which could overlap with adjacent test-related skills in the same repo.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing, 2 suspicious

Warning

Total

15

/

16

Passed

Repository
vercel/eve
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.