CtrlK
BlogDocsLog inGet started
Tessl Logo

autolab-hermes-delegation

Use Hermes delegate_task cleanly in this repo for planner, reviewer, researcher, reporter, experiment-worker, and memory-keeper roles.

80

5.00x
Quality

73%

Does it follow best practices?

Impact

100%

5.00x

Average score across 3 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./projects/pre-training/.agents/skills/autolab-hermes-delegation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

83%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is concise and highly actionable, with executable templates and concrete commands throughout. Its main gap is the worker flow's lack of explicit validation/verification checkpoints before irreversible operations like submit_patch.

Suggestions

Add a validation checkpoint in the Experiment Worker Flow before submit_patch.py (e.g. confirm the benchmark metric parsed successfully and the worktree diff is single-variable).

Include a brief feedback loop for failed runs (parse-metric failure -> log failure state -> do not submit) so the destructive submit step is guarded.

Consider splitting the five role delegate_task templates into a referenced file once the body grows, to keep SKILL.md as an overview.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence: constraints, parent rules, role toolsets, and copy-paste delegate_task templates with no padding or explanation of concepts Claude already knows; every section earns its place.

5 / 5

Actionability

Provides fully executable, copy-paste-ready delegate_task(...) blocks for five roles plus concrete uv run commands in the worker flow, covering the common cases with specific toolsets and goals.

5 / 5

Workflow Clarity

The Experiment Worker Flow is a clear numbered sequence but lacks validation checkpoints before destructive/irreversible steps (running a managed experiment and submit_patch.py); per the rubric, missing validation in batch/destructive workflows caps this at 3.

3 / 5

Progressive Disclosure

Well-organized into clearly headed sections (Parent Rules, Role Defaults, per-role templates, worker flow) with no nested references; the all-inline single-file structure is appropriate for the size, though it slightly exceeds the under-50-line simple-skill exception that would allow a 5.

4 / 5

Total

17

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and highly distinct, but it lacks an explicit 'when to use' trigger clause and leans on repo-specific jargon rather than natural user keywords. Adding a 'Use when...' sentence with everyday phrasings would lift completeness and trigger-term quality.

Suggestions

Add an explicit trigger clause, e.g. 'Use when Hermes is the parent control plane and you need to delegate a role-specific task to a child agent.'

Soften jargon with natural synonyms users might say (e.g. 'delegate work', 'hand off a task') alongside 'delegate_task'.

State the core action more concretely than 'cleanly' (e.g. 'pass the full role contract in every delegate_task call').

DimensionReasoningScore

Specificity

Names the concrete tool ("Hermes delegate_task") and enumerates six specific roles (planner, reviewer, researcher, reporter, experiment-worker, memory-keeper), but describes a single action applied across roles rather than multiple distinct actions, so it falls just below the comprehensive 5 anchor.

4 / 5

Completeness

The 'what' is clear (use Hermes delegate_task for the listed roles) but there is no 'Use when...' clause or equivalent trigger guidance; per the rubric a missing explicit trigger caps completeness at 3.

3 / 5

Trigger Term Quality

Terms like "Hermes", "delegate_task", and the role names are relevant but heavily repo-specific jargon; natural user phrasings and synonyms are missing, matching the 'some relevant keywords but missing common variations' anchor.

3 / 5

Distinctiveness Conflict Risk

The description carves a clear niche (Hermes as parent control plane for this repo with named roles) with distinct triggers and minimal overlap with other skills.

5 / 5

Total

15

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

12

/

16

Passed

Repository
huggingface/context-course
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.