CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, update or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy. 创建新技能、修改和改进现有技能、衡量技能表现。当用户想要从零创建技能、更新或优化现有技能、运行评估测试技能、通过方差分析进行性能基准测试,或优化技能描述以提高触发准确率时使用。 MANDATORY TRIGGERS: create skill, new skill, improve skill, skill eval, benchmark skill, 创建技能, 新技能, 改进技能, 评估技能

65

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills2set/skill-creator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable with concrete commands, templates, and well-sequenced validation checkpoints, but the body is verbose and padded well past what Claude needs and inlines detail that belongs in reference files. Tightening prose and splitting the long sub-processes into references would lift the weaker dimensions.

Suggestions

Trim conversational padding (e.g. 'Cool? Cool.', the plumbers/grandparents anecdote, the 'billions a year in economic value' aside) and redundant restatements to bring the body comfortably under the 500-line target and raise conciseness.

Move the full Description Optimization sub-process and its JSON/HTML templates into a reference file, leaving SKILL.md as a concise overview with a clear pointer, to improve progressive_disclosure.

Add the missing agents/ files referenced in the body (grader.md, comparator.md, analyzer.md) or remove the pointers, since dangling references weaken navigation and bundle integrity.

DimensionReasoningScore

Conciseness

At ~480 lines the body is near its stated 500-line ceiling and carries substantial chatty padding ('Cool? Cool.', the plumbers/grandparents anecdote, 'we are trying to create billions a year in economic value here!') and restated guidance Claude already knows, putting it noticeably below the midpoint versus the score-3 'mostly efficient' anchor.

2 / 5

Actionability

Concrete, copy-paste-ready commands and code are given throughout — preflight invocations, `python -m scripts.aggregate_benchmark`, `nohup python .../generate_review.py`, JSON templates for evals/metadata/timing/feedback, and placeholder-replacement steps — covering the common cases fully per the score-5 anchor.

5 / 5

Workflow Clarity

The eval/iterate loop is clearly sequenced (spawn runs → draft assertions → capture timing → grade/aggregate/launch viewer → read feedback) with explicit checkpoints and feedback loops, but the long-form prose between steps and some implicit ordering introduce minor gaps versus the score-5 'explicit validation steps + checklists' anchor.

4 / 5

Progressive Disclosure

Structure exists (sections + one-level-deep pointers to references/schemas.md, assets/eval_review.html, and agents/*.md), but the body inlines a lot of detailed content that could live in references (full JSON templates, the entire Description Optimization sub-process) and the agents/ files it points to are not present in the bundle, leaving organization only partially realized against the score-4 anchor.

3 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly states what the skill does and when to use it, with an explicit trigger list that aids Claude's invocation. Trigger term coverage is broad but could add a few more informal/synonym phrasings users actually type.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Create new skills', 'modify and improve existing skills', 'run evals to test a skill', 'benchmark skill performance with variance analysis', 'optimize a skill's description' — with comprehensive coverage of the skill's surface, matching the score-5 anchor.

5 / 5

Completeness

It explicitly answers both 'what' (create/modify/improve/measure skills) and 'when' ('Use when users want to create a skill from scratch, update or optimize an existing skill, run evals...') with concrete trigger phrases, matching the score-5 anchor exactly.

5 / 5

Trigger Term Quality

Natural user terms are present ('create a skill from scratch', 'run evals', 'benchmark skill performance') plus an explicit MANDATORY TRIGGERS list including synonyms ('new skill', 'skill eval', 'benchmark skill'), but it lacks file-extension-style anchors and some informal phrasings a user might say, so it sits just below the comprehensive 5.

4 / 5

Distinctiveness Conflict Risk

The skill-creator niche is distinct (meta-skill for building/optimizing skills and their triggering) and the explicit trigger terms are unlikely to collide with content-handling skills, giving minimal conflict risk per the score-5 anchor.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
liliMozi/openhanako
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.