CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-evals

Author and run black-box benchmark cases for the Windmill AI generation modes (flow/app/script/cli/global) in ai_evals/. Use when adding or changing eval cases, or when running before/after benchmarks for AI chat / copilot changes.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

90%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An efficient, highly actionable skill body: executable commands, concrete good/bad examples, and project-specific gotchas with zero padding. The only gaps are the absence of an error-recovery loop for benchmark runs and a single unverifiable external reference.

DimensionReasoningScore

Conciseness

Lean and dense with project-specific knowledge only (workspace-limit gotcha, backend proxy requirement, .env key handling); no padding and no explanation of concepts Claude already knows. Every token earns its place.

5 / 5

Actionability

Copy-paste ready commands for install, listing models/cases, and running with the required env vars, plus concrete good/bad prompt and judge-checklist examples. Fully executable guidance covering the common cases.

5 / 5

Workflow Clarity

Clear run sequence (install, list, run with backend env vars) and well-structured authoring rules; the harness's validation ('deterministic validation, then LLM judging') is described but no operator error-recovery feedback loop is given. Not 5 for the missing validation checkpoint guidance; well above 3 because the sequences themselves are explicit.

4 / 5

Progressive Disclosure

Well-organized sections with case-format detail deferred via a clearly signaled one-level reference ('See ai_evals/README.md for the full case format'). Not 5 because no bundle files exist to verify the referenced path resolves, a minor organization gap.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities, an explicit 'Use when...' trigger clause, and a well-scoped niche. Minor room only in synonym coverage for trigger terms.

DimensionReasoningScore

Specificity

Lists several concrete actions ('Author and run black-box benchmark cases', 'running before/after benchmarks') with an enumerated domain (the five generation modes, ai_evals/). Not 5 because coverage has minor gaps (no mention of judging, results, or fixtures); clearly above the 1-2-action anchor at 3.

4 / 5

Completeness

Explicitly answers both: what ('Author and run black-box benchmark cases for the Windmill AI generation modes (flow/app/script/cli/global) in ai_evals/') and when ('Use when adding or changing eval cases, or when running before/after benchmarks for AI chat / copilot changes') with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

Natural phrases users would say are present ('adding or changing eval cases', 'running before/after benchmarks', 'AI chat / copilot changes'). Not 5 because common synonyms like 'tests', 'regression', or 'eval suite' are missing; well above the generic-keyword anchors at 2-3.

4 / 5

Distinctiveness Conflict Risk

Clear niche (Windmill AI generation benchmark cases, specific modes, specific directory) with distinct triggers and minimal overlap with any other skill category.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
windmill-labs/windmill
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.