CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

87

1.74x
Quality

82%

Does it follow best practices?

Impact

94%

1.74x

Average score across 3 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable content with an unusually well-specified multi-step eval workflow and strong feedback loops. Its weaknesses are conversational padding that inflates token cost, and progressive disclosure undermined by referenced files (agents/, scripts/, eval-viewer/) that are absent from the bundle.

Suggestions

Trim the conversational asides ('Cool? Cool.', the plumbers/grandparents anecdote, the 'billions in economic value' line) — they add tokens without adding instruction.

Ship the referenced agents/grader.md, agents/comparator.md, agents/analyzer.md, scripts/*, and eval-viewer/generate_review.py in the bundle, or restructure the skill so SKILL.md does not depend on files it does not include.

Move the description-optimization walkthrough and the 'What the user sees in the viewer' section into a reference file to give the body more headroom under the 500-line limit.

DimensionReasoningScore

Conciseness

The operational core is dense and useful, but several sections are padded chattiness that adds no instruction: 'Cool? Cool.', the plumbers/grandparents anecdote in 'Communicating with the user', 'we are trying to create billions a year in economic value here!', and 'Sorry in advance but I'm gonna go all caps here'. These fit 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than the severely padded 1-2 anchors.

3 / 5

Actionability

Copy-paste ready throughout: exact subagent spawn prompts with output paths, JSON templates for evals.json / eval_metadata.json / timing.json / feedback.json, exact commands ('python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>', the nohup generate_review.py invocation, 'python -m scripts.run_loop ... --max-iterations 5'). Placeholders are clearly marked and the common cases are covered.

5 / 5

Workflow Clarity

The eval loop is explicitly sequenced (spawn all runs in the same turn → draft assertions while runs are in progress → capture timing per notification → grade → aggregate → analyst pass → launch viewer → read feedback) with validation checkpoints and feedback loops ('Kill the viewer server when you're done', iteration termination criteria, baseline selection rules, and per-environment fallbacks for Claude.ai/Cowork).

5 / 5

Progressive Disclosure

References are clearly signaled and one level deep where they exist ('See references/schemas.md for the full schema', 'Read agents/grader.md ... Read them when you need to spawn the relevant subagent'), but the actual bundle contains only references/schemas.md and assets/eval_review.html — the referenced agents/grader.md, agents/comparator.md, agents/analyzer.md, scripts/*, and eval-viewer/generate_review.py are missing. The 485-line body also inlines material (viewer UI description, full description-optimization walkthrough) that could live in reference files, fitting the 'some structure but could be better organized' anchor.

3 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

Strong description: third-person, concrete action list plus an explicit 'Use when...' clause covering both creation and improvement paths. Only minor gaps in capability coverage and a few missing natural synonyms keep specificity and trigger terms at 4.

Suggestions

Mention packaging/delivering the finished skill (.skill file) since it is a supported capability in the body.

Add one or two natural synonyms users might type, e.g. 'make a new skill' or 'improve SKILL.md'.

DimensionReasoningScore

Specificity

Lists several concrete actions — 'Create new skills, modify and improve existing skills, and measure skill performance', 'run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy' — with only minor coverage gaps (e.g., packaging/delivering the finished .skill file is not mentioned). Not a 5 because a couple of capabilities the skill actually supports are absent; not a 3 because far more than 1-2 concrete actions are named.

4 / 5

Completeness

Explicitly answers both: what — 'Create new skills, modify and improve existing skills, and measure skill performance'; when — 'Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals...'. This mirrors the 5 anchor's what + explicit 'Use when' with concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural user phrasings are well covered: 'create a skill from scratch', 'edit, or optimize an existing skill', 'run evals to test a skill', 'benchmark skill performance', 'optimize a skill's description', 'triggering accuracy'. A few plausible synonyms are missing ('SKILL.md', 'make/build a skill', 'skill evaluation'), keeping it below the comprehensive 5 anchor.

4 / 5

Distinctiveness Conflict Risk

A clear niche (skill authoring and evaluation tooling) with distinct triggers ('run evals to test a skill', 'benchmark skill performance with variance analysis', 'optimize a skill's description for better triggering accuracy') that no adjacent skill (code review, general evals) would claim; minimal conflict risk.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
av/harbor
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.