CtrlK
BlogDocsLog inGet started
Tessl Logo

vally-eval

Author, validate, and run Vally evaluation suites for agent skills. TRIGGERS: create eval, write eval, add eval, run eval, validate eval, vally eval, eval.yaml, add stimulus, map test to eval, migrate test to eval, eval graders, eval scoring, add eval to CI.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable reference with concrete commands, clear sections, and appropriate offloading of CI detail to a one-level reference. Could improve by tightening the framing prose and inlining a minimal eval.yaml authoring example.

Suggestions

Tighten the opening paragraph and 'Why is there a custom executor' section to remove context Claude can infer, keeping only repo-specific rationale.

Inline a minimal annotated eval.yaml example so the authoring workflow is actionable without following the external doc link.

Add an explicit 'validate before running' checkpoint inside the run-locally workflow rather than keeping validation as a standalone section.

DimensionReasoningScore

Conciseness

The body is mostly efficient with concrete commands and minimal padding, though the opening paragraph and 'Why is there a custom executor' section add context Claude could partly infer. Efficient with minor instances of over-explanation that could be trimmed.

4 / 5

Actionability

Provides mostly executable, copy-paste-ready commands (e.g. 'npm run vally validate-stimulus', 'npm run test:vally -- --plugin $PLUGIN_DIR --skill $SKILL', the npx grade command) with concrete paths. Not a 5 because the authoring section defers schema details to an external doc rather than giving inline copy-paste examples.

4 / 5

Workflow Clarity

Each workflow (write, validate, run locally, CI, re-grade) is clearly sequenced with commands, and the re-grade section includes an explicit feedback loop ('keep tuning... until results meet expectations'). Not a 5 because validation is a separate section rather than an inline checkpoint within the run workflow.

4 / 5

Progressive Disclosure

Well-organized sections with a clearly signaled one-level-deep local reference ([ci-test](./references/ci-test.md), a real file) for CI details and external links to official docs. Not a 5 because some source-file references use deeply nested relative paths (e.g. ../../../tests/vally/tag-helpers.ts) that are slightly buried.

4 / 5

Total

16

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concise what-statement, explicit and comprehensive TRIGGERS clause, and a clearly distinct niche. The only minor weakness is that the capability verbs are high-level rather than finely enumerated.

DimensionReasoningScore

Specificity

Lists several concrete actions ('Author, validate, and run Vally evaluation suites') targeting a specific domain, though the capability verbs stay high-level rather than enumerating sub-tasks. Not a 5 because coverage of concrete actions is not comprehensive, but clearly above the 1-2-action midpoint.

4 / 5

Completeness

Clearly states what ('Author, validate, and run Vally evaluation suites for agent skills') and explicitly provides when via a dedicated TRIGGERS clause with concrete trigger phrases. Explicit trigger guidance is present, so the missing-trigger cap does not apply.

5 / 5

Trigger Term Quality

The TRIGGERS clause gives comprehensive coverage of natural phrases users would say, including synonyms (create/write/add eval), a file extension (eval.yaml), and many concrete task phrasings. Matches the anchor for comprehensive coverage including synonyms and file extensions.

5 / 5

Distinctiveness Conflict Risk

Targets a clear niche (Vally eval suites for agent skills) with distinct, tool-specific triggers ('vally eval', 'eval.yaml'), giving minimal overlap risk with other skills. Third-person imperative voice is used correctly.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 suspicious

Warning

Total

15

/

16

Passed

Repository
microsoft/GitHub-Copilot-for-Azure
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.