CtrlK
BlogDocsLog inGet started
Tessl Logo

compare-results

Establish baseline-vs-candidate evaluation plans, delegate missing evaluations, compare validated results, and decide quantization feasibility. Use when the user asks to compare baseline vs quantized runs, explain an accuracy drop/regression, verify whether a quantized checkpoint is acceptable, or compare NEL/MLflow evaluation outputs. Do NOT use for generic single-model evaluation without comparison intent (use evaluation), live NEL status/debugging (use launching-evals), or generic MLflow browsing without a comparison goal (use accessing-mlflow).

80

Quality

100%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered instruction skill: lean and assumption-respecting, with concrete pointers, an explicitly gated multi-step workflow, and clean section organization that defers detail to referenced companion files. The only nit is a version-specific model name that could eventually date the guidance.

DimensionReasoningScore

Conciseness

Dense, domain-specific procedural guidance that assumes Claude's competence — it does not re-explain quantization, MLflow, or evaluation basics, and every line is task-specific. The single 'Kimi-K2.6' illustrative model name is version-specific but earns its place as a concrete example.

3 / 3

Actionability

Instruction-only yet highly actionable: concrete file paths and section names ('Read and perform .agents/skills/evaluation/references/run-validation.md External Baseline Sanity Check', 'use the canonical score field from …/recipes/tasks/<task>.md Score Extraction section'), specific settings to match, and an explicit report field list.

3 / 3

Workflow Clarity

An 8-step sequenced workflow with explicit validation gates and feedback loops — Step 4 blocks comparison until a run passes verification, Step 6 blocks a success verdict on a failed baseline ('correct and rerun it first'), and the checklist prescribes rerun-or-label recovery when items differ.

3 / 3

Progressive Disclosure

No own bundle files, but the body is organized into three clear sections and points to detailed companion-skill materials at one level deep with full paths and section names rather than duplicating them, keeping the overview navigable.

3 / 3

Total

12

/

12

Passed

Description

100%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, third-person description that names concrete actions, gives natural trigger terms, answers both what and when, and explicitly disambiguates from neighboring skills via negative triggers. It is a strong, low-conflict-risk description.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions — 'Establish baseline-vs-candidate evaluation plans, delegate missing evaluations, compare validated results, and decide quantization feasibility' — rather than vague domain naming.

3 / 3

Completeness

Explicitly answers both what (the four actions) and when ('Use when the user asks to compare baseline vs quantized runs…'), with an explicit 'Use when' trigger clause.

3 / 3

Trigger Term Quality

Covers natural user phrasings — 'compare baseline vs quantized runs, explain an accuracy drop/regression, verify whether a quantized checkpoint is acceptable, or compare NEL/MLflow evaluation outputs' — the terms a user would actually say.

3 / 3

Distinctiveness Conflict Risk

Clear niche plus explicit negative disambiguation ('Do NOT use for… use evaluation / launching-evals / accessing-mlflow') makes conflict with sibling skills unlikely.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
NVIDIA/Model-Optimizer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.