CtrlK
BlogDocsLog inGet started
Tessl Logo

waza-runner

Run evaluations on Agent Skills to measure their effectiveness. USE FOR: "run skill evals", "evaluate my skill", "test skill quality", "check skill triggers", "skill compliance check", "measure skill performance", "run evals on [skill-name]", "grade skill execution". DO NOT USE FOR: writing skills (use skill-authoring), improving frontmatter (use sensei), or general testing unrelated to skills.

59

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./waza-runner/SKILL.md
SKILL.md
Quality
Evals
Security

Skill Eval Runner

Evaluate Agent Skills like you evaluate AI Agents

This skill runs evaluations on other skills to measure their effectiveness using the same patterns that power AI agent evaluations.

When to Use

  • Running quality evaluations on a skill
  • Testing if a skill triggers on correct prompts
  • Measuring skill behavior quality
  • Generating eval reports for CI/CD

Commands

Run Evals

Run evals on <skill-name>

Initialize Eval Suite

Create evals for <skill-name>

Generate Report

Generate eval report for <skill-name>

Workflow

  1. Check for Eval Suite: Look for eval.yaml in the skill directory
  2. Load Tasks: Parse task definitions from tasks/*.yaml
  3. Execute: Run each task through the configured graders
  4. Report: Output results in JSON or Markdown format

Metrics Measured

MetricDescriptionDefault Threshold
Task CompletionDid the skill accomplish the goal?80%
Trigger AccuracyWas skill invoked on correct prompts?90%
Behavior QualityTool calls, efficiency, reasoning70%

Grader Types

  • Code Graders: Deterministic assertions, regex matching
  • LLM Graders: Model-as-judge with configurable rubrics
  • Human Graders: Manual review workflow

Example Usage

Running Evals

# From CLI
waza run ./my-skill/eval.yaml

# Output to file
waza run ./my-skill/eval.yaml -o results.json

Interpreting Results

{
  "summary": {
    "pass_rate": 0.85,
    "composite_score": 0.82
  },
  "metrics": {
    "task_completion": { "score": 0.9, "passed": true },
    "trigger_accuracy": { "score": 0.95, "passed": true }
  }
}

References

  • Eval Specification - Full eval.yaml schema
  • Writing Tasks - Task definition guide
  • Grader Reference - Available graders
Repository
microsoft/waza
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.