CtrlK
BlogDocsLog inGet started
Tessl Logo

evals-router

Use when evaluating LLM or RAG outputs: audit eval coverage, analyze failed traces, write binary judge prompts, validate judges against labels, generate targeted synthetic cases, evaluate retrieval quality, or plan review tooling. Do not use for ordinary software test implementation.

76

Quality

95%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary router-style skill body: lean, concrete, and fully sequenced with explicit validation gates, fail-closed rules, and exact commands. The only structural weakness is bundle navigation — a portion of the references/ tree is not reachable or signaled from SKILL.md itself.

Suggestions

Add one-line entries to the References section for the remaining bundle data files (contract.yaml, knowledge-demand.yaml, task-profile.json, eval-scenarios.json) and the evals/ and scorer-calibration/ subdirectories, or state explicitly that they are consumed by the ./bin/ask tooling rather than read directly.

Note in the References section that knowledge-capsule-routing.md is the index for all capsule files, so readers know the deeper reference tree is intentional and one hop away.

DimensionReasoningScore

Conciseness

The body is a lean router: no concept explanations Claude already knows, no padding, and dense imperative sections ('Prefer deterministic file, schema, regex, command, or artifact checks over LLM judges'). Every section carries routing or proof semantics that Claude could not infer. The only trim candidate is the near-duplicate failure-classification lists in Workflow step 5 and Failure Mode, but they serve distinct purposes (result classification vs blocking ownership), so this does not drop it to anchor 4's 'minor instances of over-explanation'.

5 / 5

Actionability

Guidance is fully executable for an instruction/router skill: exact commands with flags ('./bin/ask skills package verify <skill-path> --json --robot', 'scorer-calibration', 'external-review'), concrete artifact shapes (claim-to-case map, sentence-support map, lane tables with named columns), and explicit pass conditions for every route check ('pass when every claim maps to a case or named gap'). Anchor 5's copy-paste-ready commands and specific coverage of common cases is met; anchor 4 would require missing key details, and none are found.

5 / 5

Workflow Clarity

The 10-step Workflow is dependency-ordered with explicit validation checkpoints ('Run the deterministic proof closest to the changed claim', 'Stop dependent downstream gates at the first failed prerequisite') and error-recovery feedback loops ('patch only the responsible prompt, case, judge... and rerun the same check before widening'; repair followed by a 'sibling-pattern probe'). The route-check list is a per-lane checklist with pass criteria, matching anchor 5's sequence + validation + feedback loops + checklist pattern; anchor 4 would have minor validation gaps, which are absent.

5 / 5

Progressive Disclosure

The References section is well-signaled with one-line descriptors and every referenced path exists one level deep; knowledge-capsule-routing.md explicitly serves as a first-party index into the deeper capsule files with load_when conditions, preventing hidden second-order guidance. However, scored against the actual bundle, several files and subdirectories (contract.yaml, knowledge-demand.yaml, task-profile.json, eval-scenarios.json, references/evals/, references/scorer-calibration/) are not discoverable from SKILL.md, and the bulk of the bundle sits two hops behind the routing index — minor organization gaps that fit anchor 4 rather than the fully clean split of anchor 5.

4 / 5

Total

19

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, comprehensive action list with an explicit 'Use when' trigger clause and an explicit negative boundary that separates it from ordinary testing skills. The only weakness is that a few natural synonym phrases (judge calibration, hallucination/faithfulness checks) are absent from the trigger terms.

Suggestions

Add common synonym trigger phrases users actually say, such as 'judge calibration', 'hallucination or faithfulness checks', or 'golden/broken labels', to broaden natural-term coverage.

Consider naming the primary artifacts users ask for by their colloquial names (e.g., 'evals', 'traces') to widen the trigger vocabulary.

DimensionReasoningScore

Specificity

The description lists seven concrete actions — 'audit eval coverage', 'analyze failed traces', 'write binary judge prompts', 'validate judges against labels', 'generate targeted synthetic cases', 'evaluate retrieval quality', 'plan review tooling' — which comprehensively cover the skill's domain. This matches the anchor for multiple specific concrete actions with comprehensive coverage; a score of 4 would require noticeable gaps in the action list, and none are apparent.

5 / 5

Completeness

It explicitly answers both questions: the action list states what the skill does, and 'Use when evaluating LLM or RAG outputs: ...' provides concrete trigger phrases, plus an explicit negative boundary ('Do not use for ordinary software test implementation'). This mirrors the anchor-5 example structure with both what and when stated explicitly; anchor 4 would leave the 'when' less explicit or specific, which is not the case.

5 / 5

Trigger Term Quality

Natural phrases users would say are present — 'evaluating LLM or RAG outputs', 'failed traces', 'judge prompts', 'labels', 'synthetic cases', 'retrieval quality' — giving good keyword coverage. A few natural synonyms users commonly say are missing (e.g., 'calibration', 'hallucination/faithfulness checks', 'golden labels'), so it sits between good (4) and comprehensive (5) coverage, noticeably above the midpoint but not exhaustive.

4 / 5

Distinctiveness Conflict Risk

The niche is clear (LLM/RAG evaluation, not general software testing) and the explicit exclusion 'Do not use for ordinary software test implementation' sharply reduces overlap with generic testing skills. Clear niche with distinct triggers and minimal conflict risk matches anchor 5; anchor 4 would require residual overlap with closely related skills that the negative boundary already forecloses.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
jscraik/Agent-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.