CtrlK
BlogDocsLog inGet started
Tessl Logo

launching-evals

Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher. Covers running evaluations, checking status and live progress, debugging failed runs, exporting artifacts and logs, and analyzing results. ALWAYS triggers on mentions of running evaluations, checking progress, debugging failed evals, analyzing or analysing runs or results, run directories or artifact paths on clusters, Slurm job issues, invocation IDs, or inspecting logs (client logs, server logs, SSH to cluster, tail logs, grep logs). Do NOT use for creating or modifying evaluation configs.

80

Quality

100%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured operational skill: executable Quick Reference, an ordered workflow with a mandatory monitoring checkpoint and failure branch, terse non-obvious Key Facts, and clean one-level-deep references to verified bundle files. Matches the good-overall-example pattern closely.

DimensionReasoningScore

Conciseness

Lean body that assumes Claude's competence; every section earns its place with domain knowledge Claude lacks (PPP, Slurm job pairs, HF cache requirement, payload_modifier interceptor) rather than restating known concepts.

3 / 3

Actionability

Quick Reference provides fully executable, copy-paste-ready CLI commands with appropriately parameterized placeholders (<path.yaml>, <invocation_id>) plus concrete rsync/SSH recipes for artifact retrieval.

3 / 3

Workflow Clarity

Four steps sequenced IN ORDER with an explicit validation checkpoint ('MANDATORY after every nel run': poll status repeatedly until SUCCESS/FAILED) and a feedback loop (FAILED → debug-failed-runs.md), satisfying the score-3 anchor.

3 / 3

Progressive Disclosure

Body is a concise overview with well-signaled one-level-deep references to real files (run-evaluation.md, check-progress.md, analyze-results.md, debug-failed-runs.md, benchmarks/), keeping detail out of SKILL.md while aiding navigation.

3 / 3

Total

12

/

12

Passed

Description

100%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: dense list of concrete capabilities, comprehensive natural trigger terms, explicit what/when guidance, and a negative scope boundary that prevents conflicts. Third-person voice throughout with no padding.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('Run, monitor, analyze, and debug LLM evaluations' plus 'checking status and live progress, debugging failed runs, exporting artifacts and logs, and analyzing results'), matching the score-3 anchor.

3 / 3

Completeness

Explicitly answers both what ('Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher') and when ('ALWAYS triggers on mentions of...'), with a negative boundary ('Do NOT use for creating or modifying evaluation configs').

3 / 3

Trigger Term Quality

Covers natural phrasings a user would say — 'running evaluations, checking progress, debugging failed evals, analyzing or analysing runs or results, run directories or artifact paths on clusters, Slurm job issues, invocation IDs, inspecting logs (client logs, server logs, SSH to cluster, tail logs, grep logs)'.

3 / 3

Distinctiveness Conflict Risk

Clear niche (nemo-evaluator-launcher LLM evals) with distinct triggers and an explicit exclusion ('Do NOT use for creating or modifying evaluation configs') that separates it from a config-authoring skill, making conflicts unlikely.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
NVIDIA/Model-Optimizer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.