CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluating-code-models

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured body with copy-paste commands, expected outputs, and clear workflow sequencing. Its main weakness is progressive disclosure: the three reference files in the bundle are orphaned (never referenced from SKILL.md), while troubleshooting and benchmark-detail content that belongs in them is inlined in the body instead.

Suggestions

Replace the inlined 'Common Issues' section with pointers like '**Troubleshooting**: See [issues.md](references/issues.md)' — the bundle already contains that file, but the body never references it.

Move the 'Supported Benchmarks' table detail to references/benchmarks.md and keep only the 3-4 most common benchmarks inline, linking the rest.

Link references/custom-tasks.md from the workflows (e.g., from Workflow 1 Step 1 or a 'Custom benchmarks' note) so the bundle's custom-task guidance is discoverable.

DimensionReasoningScore

Conciseness

The body is dominated by executable commands, tables, and expected output with almost no explanation of concepts Claude already knows, but the per-workflow checkbox lists duplicate the 'Step N' headers that immediately follow, and several accelerate launch blocks repeat with only minor flag differences — minor trimming opportunities that keep it below anchor 5.

4 / 5

Actionability

Every workflow gives copy-paste-ready commands (install, evaluate, Docker-run, bash comparison loop, pandas results script), the expected results JSON is shown verbatim, and a full command-reference table covers flag defaults — fully executable guidance covering the common cases.

5 / 5

Workflow Clarity

Workflows are clearly sequenced with checklists, and the generate-on-host / evaluate-in-Docker split is an explicit safety checkpoint, but there are no verification steps before comparing results (e.g., confirm n_samples and task names match across runs), and the checklists are cosmetic rather than true validation gates.

4 / 5

Progressive Disclosure

The body itself is well-sectioned, but it never links to the three bundle reference files (references/benchmarks.md, custom-tasks.md, issues.md), and it inlines a 'Common Issues' section and benchmark tables that duplicate content belonging in those files — the anchor-3 pattern of references present but not signaled with content that should be separate kept inline.

3 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states a specific capability set (named benchmarks, pass@k metric) and includes an explicit 'Use when' clause with natural trigger phrases. The only blemish is the trailing credibility sentence ("Industry standard from BigCode Project used by HuggingFace leaderboards"), which is mild padding that adds no trigger value.

Suggestions

Drop the credibility sentence ('Industry standard from BigCode Project used by HuggingFace leaderboards') or fold 'BigCode' into the trigger clause, since it adds no capability or trigger information.

Add one or two action verbs beyond 'evaluates' (e.g., 'compare code models, measure functional correctness with pass@k') to broaden the action coverage toward comprehensive.

Include common user synonyms such as 'code evaluation' or 'coding benchmarks' in the 'Use when' clause to catch slightly different phrasings.

DimensionReasoningScore

Specificity

"Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics" names concrete benchmarks and a specific metric, giving scope beyond anchor 3's minimal actions, but the description hinges on a single verb (evaluates) rather than several distinct actions, so it falls short of anchor 5.

4 / 5

Completeness

It explicitly answers what ("Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics") and when ("Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality") with concrete trigger phrases, matching anchor 5 exactly.

5 / 5

Trigger Term Quality

"benchmarking code models", "comparing coding abilities", "testing multi-language support", "measuring code generation quality" are natural phrases a user would say, plus named benchmarks like HumanEval; a few common synonyms (e.g., "code evaluation", "coding benchmarks") are missing, which keeps it below anchor 5.

4 / 5

Distinctiveness Conflict Risk

Naming HumanEval, MBPP, MultiPL-E, and pass@k carves out a clear code-generation benchmarking niche with minimal overlap risk; a user asking for these triggers would unambiguously reach this skill over general LLM-evaluation skills.

5 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.