CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-paper-template

Structures Benchmark and Evaluation papers using the five-pillar framework (Research Gap, Construction Pipeline, Evaluation Framework, Empirical Findings, optional Companion Method). Returns a completeness audit, a six-part Introduction logic chain, a Section 2-7 skeleton, and a pre-submission checklist. Use when writing a benchmark paper, structuring a benchmark paper, checking whether a benchmark idea is substantive, drafting a benchmark Introduction, or planning the data-construction pipeline or experiments.

80

Quality

100%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-constructed skill body: dense and competence-assuming, with a copy-paste prompt template and concrete exemplars making it highly actionable, a clearly sequenced four-step workflow with a pre-submission verification gate, and clean one-level-deep progressive disclosure to real reference files.

DimensionReasoningScore

Conciseness

The body is information-dense and never explains concepts Claude already knows; every section (comparison table, five pillars, flowchart, skeleton, prompt template, exemplars) serves a distinct purpose, and the apparent duplication (pillars restated in the self-contained prompt block; a trailing reference index) is intentional structural design rather than padding.

3 / 3

Actionability

As an instruction-only skill it provides a copy-paste prompt template with concrete fillable slots, a structured four-step output spec, named exemplars (StatQA, nvBench 2.0, VisJudge-Bench), and specific per-section guidance — concrete and actionable throughout.

3 / 3

Workflow Clarity

The four-step workflow (completeness audit → Introduction chain → section outline → pre-submission checklist walk) is clearly sequenced, and Step 4's "Report any Critical or Major items that are unresolved" is an explicit verification checkpoint backed by a severity-classified checklist.

3 / 3

Progressive Disclosure

The body is a concise overview that splits stage-specific depth into eight one-level-deep reference files, each clearly signaled with a path and one-line description; all eight referenced paths map to real files in references/.

3 / 3

Total

12

/

12

Passed

Description

100%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities, provides an explicit "Use when" trigger clause with natural phrasing, answers both what and when, and occupies a distinct niche unlikely to conflict with other skills.

DimensionReasoningScore

Specificity

It lists multiple concrete outputs ("completeness audit", "six-part Introduction logic chain", "Section 2-7 skeleton", "pre-submission checklist") rather than vague actions, matching the score-3 anchor.

3 / 3

Completeness

It explicitly answers both what ("Structures Benchmark and Evaluation papers... Returns a completeness audit...") and when (explicit "Use when..." clause with concrete triggers), satisfying the score-3 anchor.

3 / 3

Trigger Term Quality

The "Use when" clause covers natural phrases a user would say ("writing a benchmark paper", "drafting a benchmark Introduction", "planning the data-construction pipeline", "checking whether a benchmark idea is substantive"), giving good keyword coverage.

3 / 3

Distinctiveness Conflict Risk

It carves a clear niche (benchmark/evaluation papers) with distinct triggers unlikely to fire for unrelated skills; the body further disambiguates against the sibling tech-paper-template skill.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUSTDial/Supervisor-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.