CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-paper-template

Structures Benchmark and Evaluation papers using the five-pillar framework (Research Gap, Construction Pipeline, Evaluation Framework, Empirical Findings, optional Companion Method). Returns a completeness audit, a six-part Introduction logic chain, a Section 2-7 skeleton, and a pre-submission checklist. Use when writing a benchmark paper, structuring a benchmark paper, checking whether a benchmark idea is substantive, drafting a benchmark Introduction, or planning the data-construction pipeline or experiments.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body that offloads depth to eight real reference files and provides a copy-paste prompt template plus concrete exemplars. The main gaps are minor verbosity in the framing and an only-implied validation feedback loop.

Suggestions

Tighten the conceptual opening ('A Benchmark paper does not win by proposing a new algorithm. It wins by...') and collapse near-duplicate trigger phrasings to reduce token overhead without losing the concrete outputs.

Make the pre-submission self-check an explicit feedback loop: after walking the checklist, add 'revise the affected section and re-run the checklist until no Critical/Major items remain' so validation->fix->retry is stated, not implied.

Note that references/orchestrator-notes.md is optional historical context and inlines a pointer to instantiation-template.md; flag it as non-essential so the reference graph stays cleanly one level deep.

DimensionReasoningScore

Conciseness

The body is dense with specialized structural guidance Claude would not already know and avoids explaining basic concepts, but the conceptual opening framing and some near-duplicate trigger phrasing could be trimmed. It sits above the 'mostly efficient' anchor yet below fully lean.

4 / 5

Actionability

For an instruction-only skill it is highly actionable: a copy-paste-ready prompt template, concrete per-pillar rules (e.g., 'cite at least three prior benchmarks... no more than three'), a section skeleton naming figures/tables, and specific named exemplars covering the common cases.

5 / 5

Workflow Clarity

The four-step prompt workflow (audit -> Introduction -> outline -> pre-submission self-check) is clearly sequenced and ends in an explicit checklist validation step, but the fix-and-recheck feedback loop is only implied ('Report any Critical or Major items that are unresolved') rather than stated as an explicit retry cycle.

4 / 5

Progressive Disclosure

SKILL.md is a clean overview with well-signaled one-level-deep 'Deep dive' pointers and a dedicated References section; all eight referenced files exist in references/ and the only inter-reference hop (orchestrator-notes -> instantiation-template) is contextual rather than nesting real content.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that answers what and when with concrete trigger phrases and enumerates multiple distinct outputs. It occupies a clear niche and avoids vague fluff or over-claims.

DimensionReasoningScore

Specificity

Names the domain and enumerates multiple concrete outputs — 'completeness audit, a six-part Introduction logic chain, a Section 2-7 skeleton, and a pre-submission checklist' — for comprehensive, non-generic coverage.

5 / 5

Completeness

Explicitly answers both what (the five-pillar framework and four returned artifacts) and when (a concrete 'Use when...' clause), matching the anchor for clearly answering both with concrete trigger phrases.

5 / 5

Trigger Term Quality

The 'Use when...' clause packs natural, varied trigger phrases a researcher would say — 'writing a benchmark paper', 'structuring a benchmark paper', 'drafting a benchmark Introduction', 'planning the data-construction pipeline or experiments' — with synonyms across phrasings.

5 / 5

Distinctiveness Conflict Risk

The 'Benchmark and Evaluation papers' niche with the five-pillar framing is specific and every trigger is scoped to benchmark papers, giving it a clear niche with minimal conflict risk against general paper-writing skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUSTDial/Supervisor-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.