CtrlK
BlogDocsLog inGet started
Tessl Logo

eagle3-validate

Validate that an EAGLE3 pipeline run completed successfully end-to-end. Checks all 4 steps produced expected artifacts, verifies acceptance rate meets threshold (>= 2.1), and produces a summary report. Use when user wants to verify a pipeline run or check benchmark results.

74

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable validation workflow with concrete commands, exact log signatures, explicit validation gates, and a ready-to-use report template. It respects the context budget and sequences a multi-step pipeline check with clear feedback loops.

DimensionReasoningScore

Conciseness

Lean and token-efficient: jumps straight into commands, tables, and exact log signatures without explaining what EAGLE3 is or how logging works, assuming Claude's competence throughout.

3 / 3

Actionability

Fully executable guidance — concrete bash ('ls -td experiments/cicd/cicd_*', 'find ... tail -50'), exact log strings to grep ('Average Acceptance Length {... ratio: Z.ZZ}', 'DUE TO TIME LIMIT'), explicit artifact paths, and a copy-paste report template.

3 / 3

Workflow Clarity

Clear 6-step sequence with explicit validation checkpoints (pass/fail/timeout signals in Step 1, a PASS/FAIL threshold table in Step 3) and a feedback loop ('If any task failed, suggest running /eagle3-triage').

3 / 3

Progressive Disclosure

Self-contained single-file skill with well-organized numbered sections, tables, and templates; no nested or deeply chained references, appropriate for a focused validation checklist.

3 / 3

Total

12

/

12

Passed

Description

85%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concrete description that clearly states both capability and trigger conditions, scoped to a distinct niche. The main weakness is trigger-term breadth — it covers the obvious phrasings but omits common synonyms a user might naturally say.

Suggestions

Broaden the 'Use when' clause with common natural phrasings, e.g. 'validate a run', 'did the run pass', or 'check acceptance rate / AR'.

Consider naming the key metric users reference ('acceptance rate') directly in the trigger rather than only 'benchmark results'.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Checks all 4 steps produced expected artifacts, verifies acceptance rate meets threshold (>= 2.1), and produces a summary report' — rather than vague language.

3 / 3

Completeness

Explicitly answers both what it does ('Validate... Checks all 4 steps... verifies acceptance rate... produces a summary report') and when to use it via a clear 'Use when' clause.

3 / 3

Trigger Term Quality

Includes relevant natural triggers ('verify a pipeline run', 'check benchmark results') but misses common variations users might say ('validate', 'did my run pass', 'acceptance rate').

2 / 3

Distinctiveness Conflict Risk

Targets a specific niche (EAGLE3 pipeline, 4 steps, AR >= 2.1) with distinct triggers unlikely to fire for unrelated skills.

3 / 3

Total

11

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
NVIDIA/Model-Optimizer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.