CtrlK
BlogDocsLog inGet started
Tessl Logo

waza-runner

Run evaluations on Agent Skills to measure their effectiveness. USE FOR: "run skill evals", "evaluate my skill", "test skill quality", "check skill triggers", "skill compliance check", "measure skill performance", "run evals on [skill-name]", "grade skill execution". DO NOT USE FOR: writing skills (use skill-authoring), improving frontmatter (use sensei), or general testing unrelated to skills.

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./waza-runner/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is lean and reasonably actionable with real CLI examples and compact tables, but it falls short on operational rigor: the workflow lacks validation checkpoints, and most of its reference links point to files that are missing from the bundle. Fixing the broken references and adding error-handling steps would lift both weak dimensions.

Suggestions

Create the missing references/WRITING-TASKS.md and references/GRADERS.md files, or remove their links from the References section — two of three linked files are dead.

Add validation checkpoints to the Workflow, e.g., what to do when eval.yaml is missing or a task file fails to parse, before proceeding to execution and reporting.

Ground the abstract Workflow steps with the actual commands (e.g., the waza CLI invocation for each stage) instead of leaving them as high-level descriptions.

DimensionReasoningScore

Conciseness

The body is table-driven and terse ('Code Graders: Deterministic assertions, regex matching') with no explanations of concepts Claude already knows. Minor redundancy remains — the 'When to Use' bullets largely duplicate the description's USE FOR list, and the tagline/intro sentence adds little — so it falls just short of the every-token-earns-its-place anchor.

4 / 5

Actionability

It provides executable commands ('waza run ./my-skill/eval.yaml -o results.json') and a concrete sample results JSON. However, the Workflow steps are abstract ('Parse task definitions from tasks/*.yaml', 'Run each task through the configured graders') without the corresponding commands, leaving minor gaps typical of the mostly-executable anchor.

4 / 5

Workflow Clarity

The four-step sequence (Check for Eval Suite, Load Tasks, Execute, Report) is clearly ordered, but validation checkpoints are entirely absent — no handling of a missing or malformed eval.yaml, no verify-before-report loop. This matches the sequence-present-but-checkpoints-missing anchor rather than the minor-gaps level above.

3 / 5

Progressive Disclosure

Sections are well organized and references are one level deep and clearly labeled, but 2 of the 3 referenced files (references/WRITING-TASKS.md and references/GRADERS.md) do not exist in the bundle — only EVAL-SPEC.md is present. Dead reference links are a substantial navigation failure, keeping this at the some-structure-but-could-be-better-organized anchor.

3 / 5

Total

14

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit, natural trigger phrases and clear boundary guidance against adjacent skills. Its main weakness is specificity: it states only a generic purpose rather than the concrete evaluation capabilities the skill actually provides.

Suggestions

Replace 'measure their effectiveness' with the concrete metrics the skill evaluates, e.g., 'Scores task completion, trigger accuracy, and behavior quality against configurable thresholds.'

Mention the output artifacts (JSON/Markdown reports for CI/CD) in the description so users know what the evaluation produces.

DimensionReasoningScore

Specificity

The description names the domain ('Agent Skills') but offers only one generic action phrase ('Run evaluations... to measure their effectiveness'), omitting the concrete capabilities the body documents (task completion, trigger accuracy, behavior quality metrics). It matches the anchor for naming a domain with 1-2 non-comprehensive actions, not the level-above anchor requiring several specific actions.

3 / 5

Completeness

It explicitly answers both what ('Run evaluations on Agent Skills to measure their effectiveness') and when (a quoted USE FOR trigger list plus DO NOT USE FOR exclusions). This matches the top anchor: both what and when with concrete trigger phrases.

5 / 5

Trigger Term Quality

The USE FOR list contains natural user phrasings ('run skill evals', 'evaluate my skill', 'test skill quality', 'check skill triggers', 'grade skill execution') with good variation. Coverage is good but a few natural synonyms (e.g., 'benchmark', 'assess', 'score my skill') are missing, so it does not reach the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

The DO NOT USE FOR clause explicitly disambiguates from the nearest competing skills ('writing skills (use skill-authoring)', 'improving frontmatter (use sensei)') and carves out a clear eval-runner niche with distinct triggers, matching the minimal-conflict-risk anchor.

5 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

14

/

16

Passed

Repository
microsoft/waza
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.