CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluating-llms-harness

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

72%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable skill body with concrete commands and well-structured external references, but it carries redundant checklist padding and omits validation checkpoints in its batch evaluation workflows, capping conciseness and workflow clarity.

Suggestions

Remove the redundant "Copy this checklist" boxes that restate the numbered steps already shown below them; the labeled steps are sufficient.

Add explicit verification steps to the batch model-comparison and checkpoint workflows (e.g., check the output JSON exists and has expected metrics before plotting or moving to the next model).

Tighten the workflow section intros (e.g., "Evaluate checkpoints during training.") and the repeated 'Step N:' framing to reduce token overhead.

DimensionReasoningScore

Conciseness

Mostly efficient with extensive copy-paste commands, but the per-workflow "Copy this checklist" boxes (e.g., "Benchmark Evaluation: Step 1: Choose benchmark suite") and repeated step intros add padding Claude does not need.

2 / 3

Actionability

Provides fully executable, copy-paste-ready commands with real model names, concrete flags, and an expected JSON output block — matching the highest anchor for specific executable guidance.

3 / 3

Workflow Clarity

Steps are clearly sequenced with checklists, but the batch model-comparison loop (eval_all_models.sh) and checkpoint eval lack explicit validation/verification checkpoints, which the rubric caps at 2 for batch operations.

2 / 3

Progressive Disclosure

SKILL.md is a concise overview with four well-signaled one-level-deep references under "Advanced topics" (benchmark-guide, custom-tasks, api-evaluation, distributed-eval), all of which exist as real files — matching the top anchor.

3 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that states concrete capabilities, includes an explicit "Use when..." trigger clause, and stakes out a distinct niche with named industry adopters. It is concise yet complete.

DimensionReasoningScore

Specificity

Lists multiple concrete actions and benchmarks — "Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag)" and "Supports HuggingFace, vLLM, APIs" — matching the highest anchor for multiple specific concrete actions.

3 / 3

Completeness

Explicitly answers both what (evaluates LLMs across academic benchmarks) and when via a clear "Use when..." clause, matching the top anchor for explicit triggers.

3 / 3

Trigger Term Quality

Natural user phrasing is well covered via "Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress" — phrases a user would actually say.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (academic LLM benchmarking) with distinct triggers and named adopters ("Industry standard used by EleutherAI, HuggingFace, and major labs"), making wrongful triggering unlikely.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.