CtrlK
BlogDocsLog inGet started
Tessl Logo

claude-code

Delegate coding to Claude Code CLI (features, PRs).

49

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/autonomous-ai-agents/claude-code/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable — concrete commands, realistic examples, and clearly sequenced orchestration workflows with good dialog-handling and monitoring guidance. However, it is a monolithic 750-line reference dump with no progressive disclosure: the CLI flag reference, slash-command catalog, keyboard shortcuts, and settings docs should live in reference files, and several padded asides could be trimmed.

Suggestions

Move the Complete CLI Flags Reference, slash-command tables, keyboard shortcuts, and Settings/Hooks/MCP reference sections into references/ files (e.g. cli-reference.md, interactive-reference.md), keeping SKILL.md as a concise overview with clearly signaled one-level-deep links.

Trim padded asides ('This is the cleanest integration path', the 'ultrathink' pro tip, keyboard shortcuts irrelevant to tmux orchestration) to reduce token cost without losing actionability.

Add an explicit validation checkpoint for print-mode results — e.g. check the JSON subtype field for 'success' vs 'error_max_turns'/'error_budget' before reporting results — to close the workflow's main validation gap.

DimensionReasoningScore

Conciseness

The ~750-line body inlines a full CLI flag reference, slash-command catalog, keyboard shortcuts, and settings documentation, with padded asides ('This is the cleanest integration path', the 'ultrathink' pro tip) — noticeably verbose even though most lines are dense facts rather than basic-concept explanations, so it sits between anchors 1 and 2 but closer to 2.

2 / 5

Actionability

Every section provides copy-paste-ready terminal() invocations with concrete flags, tmux send-keys/capture-pane sequences, realistic JSON output examples, and specific flag syntax — fully executable guidance covering the common delegation cases.

5 / 5

Workflow Clarity

The two orchestration modes are clearly sequenced with explicit when-to-use lists, dialog handling instructs reading the prompt before answering, and monitoring uses capture-pane with concrete TUI status indicators plus subtype-based success/error detection — but a final 'validate the result before reporting' checkpoint is implied rather than explicit, matching anchor 4's 'minor validation gaps'.

4 / 5

Progressive Disclosure

There are no bundle files at all; hundreds of lines of flag tables, slash commands, and settings reference that clearly belong in separate reference files are inlined in SKILL.md. Section headers give it more structure than anchor 2's 'no section headers' example, but the complete absence of any external references and the inlined reference material place it at 2.

2 / 5

Total

13

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names its target tool distinctly, but it is cryptic and incomplete: it lacks any explicit 'Use when' trigger guidance, offers only one generic action, and relies on terse shorthand ('(features, PRs)') instead of concrete capabilities or natural trigger phrases. It sits at the midpoint of the rubric — functional but under-signaled for skill triggering.

Suggestions

Add an explicit 'Use when...' trigger clause, e.g. 'Use when delegating coding tasks — bug fixes, feature implementation, PR review, or refactoring — to Claude Code instead of coding directly.'

Replace the cryptic '(features, PRs)' parenthetical with 2-3 concrete actions users would recognize, e.g. 'runs autonomous coding sessions, reviews PRs, executes multi-step refactors'.

Include natural trigger synonyms and variations users would say (code review, coding agent, CLI delegate, worktree) to improve trigger-term coverage.

DimensionReasoningScore

Specificity

The description names the domain ('Delegate coding to Claude Code CLI') with one concrete action, but the '(features, PRs)' parenthetical is cryptic shorthand rather than concrete capabilities, leaving it short of the 'several specific actions' of anchor 4.

3 / 5

Completeness

The 'what' is present but terse, and there is no 'Use when...' clause — the 'when' is only weakly implied by '(features, PRs)', which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Terms like 'coding', 'Claude Code', and 'PRs' are natural, but common variations users would actually say (code review, refactoring, bug fix, autonomous coding agent) are missing, matching anchor 3's 'some relevant keywords but missing common variations'.

3 / 5

Distinctiveness Conflict Risk

Explicitly naming 'Claude Code CLI' carves a clear niche distinct from sibling skills (codex, opencode, hermes-agent per related_skills), but broad triggers like 'coding' and 'PRs' still carry minor overlap risk with those close siblings, fitting anchor 4.

4 / 5

Total

13

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (755 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.