CtrlK
BlogDocsLog inGet started
Tessl Logo

arxiv

Search arXiv papers by keyword, author, category, or ID.

59

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/research/arxiv/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, highly actionable reference: executable commands throughout, a working helper script, a complete research workflow, and thoughtful edge-case guidance (ID versioning, withdrawn papers). Its main weaknesses are mild duplication between the Quick Reference, later sections, and the Notes, and inline Semantic Scholar/BibTeX content that could live in reference files.

DimensionReasoningScore

Conciseness

The body is a dense reference of API facts Claude does not know (endpoints, query syntax, sort params, rate limits) with no concept explanations, but has minor trimmable duplication: the Quick Reference table repeats commands shown in later sections, and the Notes section restates URL patterns already given. This matches anchor 4 (efficient, minor instances that could be trimmed); not 5 because the duplication is real, not absent.

4 / 5

Actionability

Every section provides copy-paste-ready curl commands and complete Python parsing snippets (e.g. the Atom XML parser and BibTeX generator), plus a real, verified helper script (scripts/search_arxiv.py) with concrete usage examples and an end-to-end workflow with exact commands. This matches anchor 5 (fully executable, covers the common cases).

5 / 5

Workflow Clarity

The "Complete Research Workflow" section sequences 7 steps with exact commands, and the withdrawn-papers section adds a checkpoint ("Always check the summary before treating a result as a valid paper"). All operations are read-only, so the destructive/batch validation cap does not apply. It matches anchor 4 (clear sequence, most checkpoints present); not 5 because there are no explicit error-recovery loops for API failures beyond stating rate limits.

4 / 5

Progressive Disclosure

The body is well-organized with headers, a Quick Reference up front, and a clearly signaled one-level-deep bundle file (scripts/search_arxiv.py, which exists and is documented with usage). Semantic Scholar and BibTeX details are inlined rather than split into references/, which is borderline but reasonable at this size. This matches anchor 4 (good structure, most content appropriately placed, minor organization gaps); not 5 because the sizable secondary-API content is not split out.

4 / 5

Total

17

/

20

Passed

Description

57%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and distinct with a clear "what", but it lacks any "when to use" trigger guidance and misses natural synonym terms (preprints, research papers, academic literature). It undersells the skill's actual capabilities, which extend well beyond search into retrieval, reading, and citation workflows.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user mentions arXiv, preprints, research papers, academic papers, or wants to find/read/cite scientific literature."

Broaden the action coverage to reflect the skill's full scope: searching, fetching paper metadata and abstracts, generating BibTeX entries, and looking up citations via Semantic Scholar.

Include natural synonyms such as "preprints", "research papers", and "scientific literature" so the description matches how users actually phrase these requests.

DimensionReasoningScore

Specificity

"Search arXiv papers by keyword, author, category, or ID" names the domain and one concrete action with four concrete parameters, but covers only search — it omits retrieval, reading, and citation workflows that the skill body actually provides. This matches anchor 3 (domain plus 1-2 concrete actions, not comprehensive); it exceeds anchor 2 (the action is specific, not generic) but falls short of anchor 4 (no list of several distinct actions).

3 / 5

Completeness

The "what" is clear (search arXiv papers by keyword, author, category, or ID) but there is no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. It is not anchor 4 because the "when" is entirely absent rather than merely imprecise.

3 / 5

Trigger Term Quality

"arXiv" and "papers" are natural terms users would say, but common variations are missing: "research papers", "preprints", "academic", "scientific literature", "abstracts". This matches anchor 3 (some relevant keywords, missing common variations or synonyms); it is not anchor 4 because the gaps go beyond "a few natural terms missing".

3 / 5

Distinctiveness Conflict Risk

"Search arXiv papers" occupies a clear niche with distinct triggers — a user mentioning arXiv would unambiguously route here, with only minimal overlap risk against a hypothetical general literature-search skill. This matches anchor 5 (clear niche, distinct triggers, minimal conflict risk) and exceeds anchor 4.

5 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.