CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

89

1.74x
Quality

85%

Does it follow best practices?

Impact

94%

1.74x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an exceptionally concrete, well-sequenced workflow with copy-paste-ready commands, JSON schemas, and validation checkpoints. Its two weaknesses are stylistic padding spread across ~485 lines and a progressive-disclosure failure: several referenced bundle files (agents/*.md, eval-viewer/generate_review.py) are absent, so the skill's own mandatory steps cannot be followed as written.

Suggestions

Ship the missing bundle files referenced by the body — agents/grader.md, agents/comparator.md, agents/analyzer.md, and eval-viewer/generate_review.py — or rewrite those sections to inline the needed guidance, since the skill's mandatory steps (grading, viewer generation) currently point at nonexistent paths.

Trim persona chatter and redundancy: remove lines like "Cool? Cool.", "Good luck!", the plumbers/npm anecdote, and the "Repeating one more time the core loop" recap section, which duplicates the intro's process list and the final summary.

Consolidate the Cowork all-caps reiteration into the existing viewer-generation step rather than repeating the instruction twice; state it once with the rationale (get outputs in front of the human before self-evaluating).

DimensionReasoningScore

Conciseness

The body is mostly substantive workflow instruction, but padded sections are scattered throughout: "Cool? Cool.", the plumbers/npm anecdote, "Good luck!", "we are trying to create billions a year in economic value here!", an all-caps Cowork reiteration, and a full recap section ("Repeating one more time the core loop here for emphasis") that repeats content already stated. Not 4: these unnecessary passages are more than minor trimmable instances; not 2: the core content is efficient and assumes Claude's competence rather than explaining known concepts.

3 / 5

Actionability

Fully executable throughout: exact subagent spawn prompts ("Execute this task: - Skill path: <path-to-skill>..."), exact bash commands ("python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>", the nohup viewer launch with VIEWER_PID capture), exact JSON structures for evals.json/eval_metadata.json/timing.json, and exact grading.json field names ("text", "passed", "evidence"). Not 4: commands and examples cover the common cases copy-paste ready with no meaningful gaps.

5 / 5

Workflow Clarity

The eval loop is a clearly sequenced 5-step process with explicit validation checkpoints: spawn with-skill AND baseline runs together, draft assertions while runs proceed, capture timing data per notification as runs complete ("Process each notification as it arrives"), grade assertions, aggregate into benchmark, analyst pass, launch viewer, read feedback, then an iteration loop with explicit exit criteria ("The user says they're happy / feedback is all empty / not making meaningful progress"). Not 4: checkpoints and feedback loops are explicit and complete, including error-recovery guidance (baseline snapshotting, --static fallback for headless environments).

5 / 5

Progressive Disclosure

Structure is good — a dedicated "Reference files" section signals when to read each external file, and references/schemas.md and assets/eval_review.html exist as referenced. However, the body repeatedly directs the model to files that do not exist in the bundle: agents/grader.md, agents/comparator.md, agents/analyzer.md ("Read agents/grader.md"... "Read agents/comparator.md and agents/analyzer.md") and eval-viewer/generate_review.py (invoked in an all-caps mandatory instruction), so navigation breaks at exactly the points the skill leans on. Not 4: broken references are a worse failure than minor organization gaps; not 2: the sections are well-organized and the references that do exist are clearly signaled rather than buried or inlined.

3 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability list, and an explicit 'Use when...' clause enumerating realistic user phrasings including near-synonyms. The only gap is minor synonym coverage in trigger terms.

DimensionReasoningScore

Specificity

"Create new skills, modify and improve existing skills, and measure skill performance" plus "run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy" lists multiple specific concrete actions with comprehensive coverage of the skill's capabilities. Not 4: coverage is comprehensive rather than having minor gaps.

5 / 5

Completeness

Explicitly answers both: what ("Create new skills, modify and improve existing skills, and measure skill performance") and when ("Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals...") with concrete trigger phrases. Not 4: the 'when' clause is fully explicit with enumerated triggers, not merely adequate.

5 / 5

Trigger Term Quality

"create a skill from scratch", "edit, or optimize an existing skill", "run evals", "benchmark skill performance", "optimize a skill's description" are natural phrases users would say. Not 5: a few common synonyms ("make/build a skill", "improve my skill's triggering", "test my skill") are not covered; not 3: keyword coverage is good, not partial.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear meta-niche (authoring/evaluating skills themselves) with distinct triggers like "run evals to test a skill" and "benchmark skill performance with variance analysis" — unlikely to be confused with domain skills. Not 4: the niche is unambiguous with minimal overlap risk.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
anthropics/claude-plugins-official
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.