CtrlK
BlogDocsLog inGet started
Tessl Logo

vally-eval

Author, validate, and run Vally evaluation suites for agent skills. TRIGGERS: create eval, write eval, add eval, run eval, validate eval, vally eval, eval.yaml, add stimulus, map test to eval, migrate test to eval, eval graders, eval scoring, add eval to CI.

69

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-organized, actionable overview: copy-paste commands, explicit validation and re-grade feedback loops, and clean progressive disclosure to two real, one-level-deep reference files. It falls slightly short of top marks on conciseness (a few rationale sentences could be trimmed) and actionability (eval-suite YAML schema and custom grader authoring rely on external documentation rather than inline examples).

DimensionReasoningScore

Conciseness

The body is dominated by repo-specific conventions and copy-paste commands Claude could not know, with no padding explaining known concepts. Minor over-explanation remains (e.g., the "Why is there a custom executor" rationale paragraph and "In most cases, you would like to use a command like this"), so it is efficient but could be trimmed slightly. Not 5 because a few sentences do not earn their tokens; not 3 because there is no unnecessary concept explanation.

4 / 5

Actionability

Concrete, executable commands appear throughout ("npm run vally validate-stimulus", "npm run test:vally -- --plugin $PLUGIN_DIR --skill $SKILL", full "npx @microsoft/vally-cli grade ... --verbose < results/<test-run-name>/results.jsonl" invocations, concrete file-layout paths). Not 5 because key authoring tasks defer to external documentation — the eval suite YAML schema ("Refer to the official documentation on the schema of the spec") and custom grader creation ("follow the examples in the official vally documentation") lack inline examples.

4 / 5

Workflow Clarity

Sections are sequenced by task (write, validate, run locally, run in CI, extend, re-grade, collect results) and include a validation checkpoint ("a script is added to validate the eval suites and report errors when it sees any violation") plus a re-grade feedback loop ("keep tuning the grader config ... and re-grade the trajectory until the results meet your expectations"), with MUST/MUST NOT constraints flagged. Not 5 because there is no explicit fix-and-re-validate loop after validation errors, and the authoring workflow itself is not laid out as ordered steps.

4 / 5

Progressive Disclosure

The body is a well-sectioned overview with clearly signaled one-level-deep references that exist in the bundle ("See [azure-fixtures](./references/azure-fixture.md) for more details", "Refer to [ci-test](./references/ci-test.md)"), and detail content (Azure fixtures, CI setup) is appropriately split into those files. Navigation is easy and nothing that belongs in references is inlined.

5 / 5

Total

17

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities in third person, provides explicit trigger guidance with comprehensive natural keywords including synonyms and a file extension, and both what and when are clearly answered. Its only weakness is a handful of generic 'eval' trigger terms that create minor overlap risk with other evaluation-related skills.

DimensionReasoningScore

Specificity

"Author, validate, and run Vally evaluation suites" lists several concrete third-person actions (author, validate, run) in a clearly named domain. Not 3 because more than 1-2 actions are named; not 5 because capabilities like re-grading trajectories and custom graders are not mentioned, leaving minor gaps in coverage.

4 / 5

Completeness

"Author, validate, and run Vally evaluation suites for agent skills" explicitly answers what the skill does, and "TRIGGERS: ..." provides explicit, concrete trigger guidance equivalent to a 'Use when...' clause. Both what and when are clearly and explicitly stated.

5 / 5

Trigger Term Quality

The TRIGGERS list covers natural phrases ("create eval", "write eval", "run eval", "validate eval"), synonyms ("map test to eval", "migrate test to eval", "add stimulus"), specialized terms ("eval graders", "eval scoring", "add eval to CI"), and a file extension ("eval.yaml"), matching comprehensive coverage.

5 / 5

Distinctiveness Conflict Risk

"vally eval", "eval.yaml", and "add stimulus" establish a clear niche with minimal conflict risk, but broadly generic terms like "create eval", "write eval", and "run eval" could overlap with other evaluation-related skills. Mostly distinct with minor overlap risk, so not 5.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 suspicious

Warning

Total

15

/

16

Passed

Repository
microsoft/GitHub-Copilot-for-Azure
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.