CtrlK
BlogDocsLog inGet started
Tessl Logo

creating-eval-scenarios

Generate evaluation scenarios for Tessl tiles to measure skill effectiveness. Creates inventory of instructions from the skill, test cases with success criteria, and validates skill coverage. Use when asked to "generate evals", "create evaluation scenarios", "test this skill", "measure skill value", or "prepare for tessl publish".

71

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Well-organized overview that practices progressive disclosure by pointing to a single detailed reference. The body is concise and actionable for structure/commands, while the multi-step workflow detail is correctly delegated to the reference.

DimensionReasoningScore

Conciseness

The body is lean and mostly assumes Claude's competence, with a few explanatory constraint bullets that could be tightened slightly.

4 / 5

Actionability

Provides a concrete output file tree and executable 'tessl eval' commands, but delegates the core how-to guidance to the reference file rather than embedding it.

4 / 5

Workflow Clarity

A clear sequence is present (prerequisites, read reference, produce files, run evals) with a feasibility checkpoint, though detailed validation loops live in the referenced file rather than the body.

4 / 5

Progressive Disclosure

Clear overview body with a single well-signaled one-level-deep reference (references/scenario-generation.md, which exists), with content appropriately split between the two.

5 / 5

Total

17

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concrete description with explicit what/when structure and natural trigger phrases. Only slight overlap risk from a couple of generic trigger terms keeps distinctiveness just below full marks.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Creates inventory of instructions from the skill, test cases with success criteria, and validates skill coverage' — giving comprehensive coverage of what the skill does.

5 / 5

Completeness

Explicitly answers both what it does and when to use it, with a 'Use when asked to ...' clause listing concrete trigger phrases.

5 / 5

Trigger Term Quality

Provides several natural quoted trigger phrases a user would actually say, including synonyms ('generate evals' / 'create evaluation scenarios') and 'prepare for tessl publish'.

5 / 5

Distinctiveness Conflict Risk

The Tessl-tile eval niche is clearly distinct, but generic triggers like 'test this skill' and 'measure skill value' carry minor overlap risk with broader testing/evaluation skills.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ericjohnolson/tessl-test
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.