CtrlK
BlogDocsLog inGet started
Tessl Logo

research-paper-writing

Write ML papers for NeurIPS/ICML/ICLR: design→submit.

58

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/research/research-paper-writing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a highly actionable, well-sequenced pipeline with strong validation feedback loops and clean one-level-deep references; its main weakness is conciseness, as the very long body explains some concepts Claude already knows and inlines material that duplicates the reference files.

Suggestions

Trim general-knowledge padding (review criteria explanations, colorblind-palette and vector-graphics advice, the ASCII pipeline diagram) to tighten the body toward the lean anchor.

Move the inline autoreason, paper-types, and reviewer-criteria summaries fully into their reference files and replace the inline text with brief pointers to reduce duplication.

Isolate time-sensitive references (model tiers like 'Haiku 3.5', 'Sonnet 4', 'S4.6' and 'NeurIPS 2025' dates) in a clearly labeled version/deprecation section so they don't penalize conciseness.

DimensionReasoningScore

Conciseness

The ~1610-line body is mostly efficient specialized guidance with executable code, but contains several padded sections and concepts Claude already knows (review criteria, colorblind palettes, the ASCII pipeline diagram) plus time-sensitive model/version references not isolated in a deprecated section, keeping it at the 'mostly efficient but could be tightened' anchor.

3 / 5

Actionability

The body is rich with copy-paste-ready executable code and commands — nohup/cron/latexmk/chktex bash, complete Python (cost tracker, citation fetcher, analysis), LaTeX snippets, JSON review templates — covering the common cases, matching the fully-executable anchor.

5 / 5

Workflow Clarity

Eight numbered phases with sequenced sub-steps and explicit validation checkpoints — the mandatory 5-step citation verification ('If ANY step fails → [CITATION NEEDED]'), pre-compilation validation, failure-recovery table, and the Step 6.3 review→fix→re-check loop — match the clear-sequence-with-feedback-loops anchor.

5 / 5

Progressive Disclosure

The body acts as an overview with clearly signaled one-level-deep references (verified to exist) and a navigation table; it falls short of 5 because the monolithic inline bulk and partial duplication of autoreason/paper-type/reviewer material already present in reference files could be offloaded.

4 / 5

Total

17

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is compact and names a clear, distinctive niche with strong natural trigger terms (the venue names), but it lacks an explicit 'Use when...' trigger clause and compresses its actions into an arrow shorthand, limiting specificity and completeness.

Suggestions

Add an explicit trigger clause, e.g. 'Use when writing, revising, or submitting an ML/AI research paper to NeurIPS, ICML, ICLR, ACL, AAAI, or COLM.'

Expand the arrow shorthand into 2-3 concrete actions to raise specificity (e.g. 'design experiments, draft and revise sections, verify citations, prepare submissions').

Include synonyms users actually say — 'research paper', 'publication', 'camera-ready' — to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

The description names the domain ('ML papers') and conveys two actions via the compressed arrow ('design→submit'), but does not list several specific concrete actions, so it sits at the 'domain + 1-2 actions, not comprehensive' anchor rather than a 4.

3 / 5

Completeness

The 'what' is clear ('Write ML papers'), but there is no explicit 'Use when...' clause or equivalent trigger guidance, and the 'when' is only weakly implied by the venue list — per the rubric, a missing Use-when clause caps completeness at 3.

3 / 5

Trigger Term Quality

'ML papers', 'NeurIPS', 'ICML', 'ICLR' are natural terms a user would say when needing this skill, giving good keyword coverage; it falls short of 5 because synonyms like 'research paper' or 'publication' and the other venues covered in the body are absent.

4 / 5

Distinctiveness Conflict Risk

'Write ML papers for NeurIPS/ICML/ICLR' pins a clear niche with minimal conflict risk, but the terse 'write papers' phrasing leaves minor overlap risk with a generic paper-writing skill, keeping it just below 5.

4 / 5

Total

14

/

20

Passed

Validation

62%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation10 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (1628 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 missing

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

10

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.