CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, update or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

89

1.90x
Quality

85%

Does it follow best practices?

Impact

95%

1.90x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

70%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a comprehensive, highly actionable skill with excellent workflow clarity and progressive disclosure. Its major weakness is extreme verbosity — conversational asides, repeated core loop summaries (3 times), communication style coaching, and filler text significantly bloat the token budget. The content would be substantially more effective at perhaps 60% of its current length.

Suggestions

Remove the three redundant restatements of the core loop (intro, end of 'Running and evaluating test cases', and the final repetition) — state it once clearly at the top.

Cut conversational filler ('Cool? Cool.', 'This task is pretty important (we are trying to create billions a year in economic value here!)', 'Sorry in advance but I'm gonna go all caps here') — these waste tokens without adding actionable guidance.

Condense the 'Communicating with the user' section to 2-3 sentences — Claude doesn't need a sociological explanation of who uses terminals now.

Tighten the 'How to think about improvements' section by removing philosophical asides and focusing on the concrete patterns (generalize, keep lean, explain why, look for repeated work) in bullet-point form.

DimensionReasoningScore

Conciseness

The skill is extremely verbose at ~500+ lines with significant conversational filler ('Cool? Cool.', 'Sorry in advance but I'm gonna go all caps here'), redundant repetitions of the core loop (stated 3 times), explanations of concepts Claude already knows (what plumbers and grandparents are doing), and lengthy asides about communication style that could be much more concise.

1 / 3

Actionability

Despite verbosity, the skill provides highly concrete, executable guidance: specific CLI commands, exact JSON schemas, file path conventions, script invocations with arguments, and detailed step-by-step procedures for spawning subagents, grading, aggregating benchmarks, and launching the viewer.

3 / 3

Workflow Clarity

The multi-step workflow is clearly sequenced with explicit validation checkpoints: capture intent → interview → write skill → run tests (with parallel baseline) → grade → aggregate → launch viewer → collect feedback → iterate. Each step has clear inputs/outputs, and there are explicit feedback loops (review → improve → rerun) with error recovery guidance.

3 / 3

Progressive Disclosure

Content is well-structured with clear references to external files (agents/grader.md, agents/comparator.md, agents/analyzer.md, references/schemas.md) that are one level deep and clearly signaled with descriptions of when to read them. The skill appropriately keeps the overview in SKILL.md while pointing to specialized subagent instructions and schema references.

3 / 3

Total

10

/

12

Passed

Description

100%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

This is a strong, well-crafted description that clearly communicates both what the skill does and when it should be used. It lists multiple concrete actions, includes natural trigger terms users would employ, and occupies a distinct niche that minimizes conflict risk with other skills. The explicit 'Use when...' clause with varied trigger scenarios is particularly effective.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions: 'create new skills', 'modify and improve existing skills', 'measure skill performance', 'run evals', 'benchmark skill performance with variance analysis', 'optimize a skill's description for better triggering accuracy'.

3 / 3

Completeness

Clearly answers both 'what' (create, modify, improve, measure skills) and 'when' with an explicit 'Use when...' clause listing specific trigger scenarios like creating from scratch, updating, running evals, benchmarking, and optimizing descriptions.

3 / 3

Trigger Term Quality

Includes strong natural keywords users would say: 'create a skill', 'update', 'optimize', 'evals', 'benchmark', 'skill performance', 'triggering accuracy', 'description'. These cover a good range of terms a user working with skills would naturally use.

3 / 3

Distinctiveness Conflict Risk

The description targets a very specific meta-domain — skill creation, modification, evaluation, and optimization — which is a clear niche unlikely to conflict with other skills. Terms like 'evals', 'variance analysis', and 'triggering accuracy' are highly distinctive.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation11 / 11 Passed

Validation for skill structure

No warnings or errors.

Repository
SnowingFox/macos-island
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.