CtrlK
BlogDocsLog inGet started
Tessl Logo

review-eval-scenario-integrity

Review a change to an eval scenario for whether its result means anything — whether the task can discriminate, and whether a judge applying its criteria reaches the same verdict twice. Use as one lens in a code review run.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./review-lenses/review-eval-scenario-integrity/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized, concise review lens with concrete reviewer procedure and clear report/don't-report decision rules. It scores top marks on progressive disclosure as a short, cleanly sectioned skill, with the only room for improvement being slightly tighter prose and more explicit step/checklist formatting.

DimensionReasoningScore

Conciseness

The body is tight prose that assumes Claude's competence and avoids explaining basics (what an eval is, what a judge is); each sentence carries review guidance. Not a 5 because a few phrases are ornate ('a score comes back looking like evidence when it is not', 'has handed over the reasoning under test') and could be trimmed; not a 3 because there is no padded over-explanation of concepts Claude already knows.

4 / 5

Actionability

Gives concrete reviewer moves ('Read the task as the agent under test sees it, with the skill absent', 'read the task beside the skill... sentence by sentence') and an explicit Reporting checklist (quote the line, name the two verdicts, state the observable). Not a 5 because the Method is procedural prose rather than copy-paste-ready steps; not a 3 because the guidance is specific and executable for an instruction-only skill.

4 / 5

Workflow Clarity

The Method is a clear ordered sequence (task first → read absent the skill → read beside the skill → take the criteria → map changed instructions to criteria), with the Threshold section acting as a report/don't-report decision checkpoint. Not a 5 because there is no explicit feedback loop or error-recovery checklist; not a 3 because the sequence and decision rules are clearly present. The destructive/batch validation cap does not apply to an analytical review skill.

4 / 5

Progressive Disclosure

The body is under 50 lines with no need for external references, and is organized into clearly signaled sections (Scope, Method, Threshold, Reporting), satisfying the simple-skill exception for a 5. Not a 4 because no content is misplaced or under-signaled; the structure is clean and navigable.

5 / 5

Total

17

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinct with an explicit 'Use as one lens' trigger, but its action framing stays abstract ('Review... for whether...') and its keywords lean technical. It earns solid marks for completeness and distinctiveness, slightly less for specificity and trigger-term naturalness.

Suggestions

Add 1-2 concrete reviewer actions to the 'what' (e.g. 'reads the task with the skill absent, then sentence-by-sentence beside it') to lift specificity toward 4-5.

Broaden trigger terms with natural phrasing a user might say (e.g. 'eval validity', 'flaky judge', 'eval that leaks the answer') alongside the technical jargon.

Make the 'when' clause slightly more concrete by naming the change context that warrants this lens (e.g. 'Use when reviewing a PR that changes an eval scenario or its judge criteria').

DimensionReasoningScore

Specificity

Names the domain (eval scenarios) and what to check for ('whether the task can discriminate', 'whether a judge... reaches the same verdict twice'), but the only action verb is 'Review', so it stops at domain-plus-1-2-concrete-concerns rather than multiple concrete actions. Not a 4 because it does not list several specific actions with minor gaps; not a 2 because it goes well beyond a bare domain mention.

3 / 5

Completeness

States the 'what' (review eval-scenario changes for discrimination and verdict stability) and an explicit 'when' ('Use as one lens in a code review run'), satisfying the trigger-guidance requirement that would otherwise cap at 3. Not a 5 because the 'when' is a single clause and could name concrete trigger phrases more explicitly; not a 3 because both what and when are clearly present.

4 / 5

Trigger Term Quality

Relevant insider keywords ('eval scenario', 'judge', 'criteria', 'verdict', 'code review run') are present, but they are technical jargon and omit common variations a user might naturally voice. Not a 4 because natural synonyms users would actually say are largely missing; not a 2 because more than one or two generic keywords appear.

3 / 5

Distinctiveness Conflict Risk

The niche (eval-scenario integrity as one review lens) is narrow and distinct, with low trigger conflict risk. Not a 5 because 'Use as one lens in a code review run' leaves minor overlap with other review-lens skills; not a 3 because it is far more specific than a broad 'works with documents'-style description.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tesslio/product-plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.