CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluating-llms-harness

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/lm-evaluation-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable content with executable commands, complete scripts, and a well-structured reference bundle. It is held back by repeated near-identical command blocks (a tightening opportunity) and the absence of validation checkpoints in batch evaluation workflows, which the rubric caps at 3.

Suggestions

Deduplicate the near-identical lm_eval invocations: define the base command once (Quick start) and show only the varying flags (--model vllm args, --num_fewshot, tensor_parallel_size) in each workflow, which would also let the Common issues section drop its repeat of Workflow 4's vLLM speedup.

Add validation checkpoints to the batch workflows: in eval_all_models.sh check the lm_eval exit status and confirm each results/<model>.json exists before continuing the loop, and in the comparison-table script skip or report models with missing/malformed result files.

Move the learning-curve plotting code, the comparison-table generation code, and the hardware requirements tables into a reference file (e.g. references/benchmark-guide.md or a new references/analysis.md) to shorten the SKILL.md body toward a lean overview.

DimensionReasoningScore

Conciseness

Mostly efficient, dense with commands and code and free of concept explanations Claude already knows, but the same lm_eval invocation is repeated near-identically ~8 times (Quick start, Workflow 1, 3, 4, and Common issues), and the Common issues 'too slow' entry duplicates Workflow 4's vLLM guidance. Not 2 (no padded prose); not 4 (the duplication goes beyond minor trimming).

3 / 5

Actionability

Fully executable, copy-paste-ready commands and complete scripts (eval_checkpoint.sh, eval_all_models.sh, plotting and comparison-table Python), with concrete example output JSON and a rendered markdown table. Covers the common cases exactly as the top anchor requires. Not 4 since no meaningful gaps remain.

5 / 5

Workflow Clarity

Workflows have clear numbered steps and checklists, but no validation checkpoints: the multi-model eval loop never checks exit status or verifies each result file exists before building the comparison table, and the checkpoint workflow doesn't verify outputs before plotting. Per the rubric cap, batch operations without validation cap workflow clarity at 3; without the cap this would be 4.

3 / 5

Progressive Disclosure

All four references/ files exist and are well-signaled one-level-deep links under 'Advanced topics', and the body is cleanly sectioned. Not 5: the ~480-line body inlines plotting code, comparison-table code, and hardware tables that could be split into references; not 3 since the split that exists is appropriate and easy to navigate.

4 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' clause, concrete benchmark names, and backend support. Its main gap is that it centers on a single action (evaluate) rather than enumerating multiple distinct capabilities, and it lacks a few common synonyms (eval, leaderboard).

DimensionReasoningScore

Specificity

Names the domain, 60+ benchmark scope with concrete examples (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag), and supported backends (HuggingFace, vLLM, APIs) — several specifics, but the action list is effectively a single verb (evaluate) rather than multiple distinct actions. Not 3 (more than 1-2 concrete specifics); not 5 (doesn't enumerate multiple distinct concrete actions).

4 / 5

Completeness

Explicitly answers what ('Evaluates LLMs across 60+ academic benchmarks...') and when with concrete trigger phrases ('Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress'), matching the top anchor exactly. Not 4 since the 'when' clause is already explicit and specific.

5 / 5

Trigger Term Quality

'benchmarking model quality, comparing models, reporting academic results, tracking training progress' are natural phrases users would say, plus named benchmarks. Missing common synonyms such as 'eval', 'evaluation', 'leaderboard', or 'run benchmarks' that anchor 5 expects.

4 / 5

Distinctiveness Conflict Risk

Clear niche in academic LLM benchmarking with distinct triggers (named benchmarks, lm-eval vocabulary), but 'comparing models' and 'tracking training progress' could overlap with general model-comparison or training-monitoring skills. Not 5 due to this minor overlap risk; clearly above 3.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.