CtrlK
BlogDocsLog inGet started
Tessl Logo

eval-corpus

Use when changing local eval capture, gold labels, diversified and tune sets, matching, replay storage, or eval CLI behavior.

56

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/eval-corpus/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is an exceptionally concrete, non-redundant reference that names exact files, functions, commands, and tests, making it highly actionable. Its weaknesses are structural: a monolithic single-heading bullet list with no sequenced validation workflow for its batch operations and no sub-section navigation.

Suggestions

Break the single bullet list into labeled sub-sections (e.g. Capture, Gold labeling, Diversified sets, Storage, Idempotency) so the document is navigable rather than a wall of bullets.

For the batch/destructive operations (capture, relabel, prune), add an explicit sequenced procedure with a validation/checkpoint step so workflow clarity can exceed the batch cap.

Split the longest run-on bullets (e.g. the gold-labeling bullet) into shorter, separately scannable rules to improve readability without adding tokens.

DimensionReasoningScore

Conciseness

The body is dense and information-rich with no padding of concepts Claude already knows — every line carries codebase-specific rules; it stops short of 5 because several bullets are very long run-on sentences that could be split or tightened for readability.

4 / 5

Actionability

It gives fully actionable, concrete guidance: exact config keys, file paths, function/owner names, CLI subcommands, and named regression tests guarding each invariant — as actionable as an instruction-only skill gets.

5 / 5

Workflow Clarity

The body states invariants and behaviors rather than a sequenced multi-step workflow, and it covers batch/destructive operations (capture, relabel, prune) without explicit validate-then-proceed checkpoints; per the batch-operation cap this holds at 3 even though guardrails like idempotency and protected cases are mentioned.

3 / 5

Progressive Disclosure

Everything is inlined under a single heading as a themed bullet list with no sub-section headers and no bundle files to reference; it has some structure (each bullet leads with its topic) but would navigate better with named sub-sections for capture, gold labeling, diversified sets, storage, and idempotency.

3 / 5

Total

15

/

20

Passed

Description

56%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a strong, specific trigger clause and a clear niche, but it omits any statement of what the skill does, leaving it incomplete. Tightening the action verb and adding a brief "what" would lift it across both completeness and specificity.

Suggestions

Add a brief "what" clause, e.g. "Documents the rules for..." or "Guides edits to local eval capture, gold labeling..." so the description answers both what and when.

Replace the single generic verb "changing" with one or two concrete actions (e.g. "capturing, labeling, replaying") to raise specificity.

Optionally include a natural synonym (e.g. "eval corpus") alongside the jargon terms so the trigger matches more phrasings users might say.

DimensionReasoningScore

Specificity

The description enumerates many concrete domain areas ("local eval capture, gold labels, diversified and tune sets, matching, replay storage, or eval CLI behavior") but names only a single generic action ("changing"), matching the anchor that names the domain with limited concrete actions.

3 / 5

Completeness

The description is exclusively a "when" clause ("Use when changing...") with no statement of what the skill does, matching the anchor-2 pattern ("Use when working with documents"); it cannot reach 3 because a clear "what" is entirely absent.

2 / 5

Trigger Term Quality

It covers the relevant area keywords a developer on this codebase would actually say (eval capture, gold labels, replay storage, eval CLI), with good breadth across the subsystem; it falls short of 5 only because the terms are jargon-heavy with no synonyms or file-extension variations.

4 / 5

Distinctiveness Conflict Risk

The trigger is tightly scoped to the internal eval-corpus subsystem with distinct, specific terms, giving it a clear niche and minimal risk of firing for an unrelated skill.

5 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
kunchenguid/no-mistakes
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.