CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluating-code-models

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

65%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with ready-to-run commands, but it loses efficiency through repeated content and fails to route detail into the existing reference files. Workflow clarity is capped by the absence of validation checkpoints on batch operations.

Suggestions

Reference the existing bundle files from the body — e.g., in Supported Benchmarks link to references/benchmarks.md and in Common Issues link to references/issues.md — instead of duplicating that content inline.

Deduplicate benchmark descriptions (currently in Quick Start, Workflow 1, and the table) into a single place, pointing elsewhere for detail.

Add explicit validation/verification checkpoints to the workflows (e.g., confirm generations saved before Docker evaluation, verify result JSON contains expected pass@k keys) so batch operations clear the workflow-clarity bar.

DimensionReasoningScore

Conciseness

Mostly efficient with executable code, but benchmark descriptions are repeated three times (Quick Start, Workflow 1, and the Supported Benchmarks table) and CLI flags appear both inline and in a Command Reference table — material that could be tightened without loss.

2 / 3

Actionability

Provides fully executable, copy-paste-ready accelerate/docker/bash/python commands with concrete flags and expected JSON output, matching the executable-code anchor.

3 / 3

Workflow Clarity

Workflows have numbered steps and checklists, but these are batch operations (model loops, code execution) with no validation/verification checkpoints, which per the rubric caps workflow clarity at 2 rather than 3.

2 / 3

Progressive Disclosure

Three reference files exist (benchmarks.md, custom-tasks.md, issues.md) but the body never signals or links to them, while overlapping benchmark and troubleshooting content is duplicated inline — content that should be separate is inline and references are present but not signaled.

2 / 3

Total

9

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: third-person, concrete, with explicit Use-when triggers and a distinct niche. Only minor weakness is the slight marketing flourish at the end, which does not meaningfully hurt the score.

DimensionReasoningScore

Specificity

Lists multiple concrete actions on named benchmarks — "Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics" — matching the multi-specific-action anchor.

3 / 3

Completeness

Explicitly answers both what ("Evaluates code generation models across...") and when ("Use when benchmarking code models, comparing coding abilities...") with a clear trigger clause.

3 / 3

Trigger Term Quality

Natural user phrasing is well covered: "benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality", terms a user would actually say.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (BigCode code-model benchmarking with pass@k) unlikely to collide with general LLM skills; the trailing "Industry standard from BigCode Project used by HuggingFace leaderboards" is mild marketing but reinforces rather than blurs the niche.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.