CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-creator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a well-sequenced, highly actionable eval-and-iterate workflow with real templates, commands, and feedback loops, but it is padded with conversational asides and redundant repetition, and several of its bundle pointers (the entire agents/ directory and eval-viewer/generate_review.py) do not exist in the shipped skill.

Suggestions

Trim conversational filler ("Cool? Cool.", the plumber anecdote, "Good luck!") and de-duplicate the viewer-generation and update-an-existing-skill instructions, which are each repeated across 2-4 sections.

Fix dangling bundle references: either ship the agents/ subagent instructions and eval-viewer/generate_review.py the body repeatedly points to, or update the paths to files that actually exist in the bundle.

Move the Claude.ai- and Cowork-specific adaptation sections into a references/ file (e.g., references/environments.md) with a one-line pointer from the main body to cut length toward the skill's own 500-line guidance.

DimensionReasoningScore

Conciseness

The ~486-line body carries several padded conversational passages — "Cool? Cool.", the plumbers-and-grandparents anecdote, "we are trying to create billions a year in economic value here!", "Good luck!" — and repeats the same instructions multiple times (the generate_review.py viewer instruction appears in the overview, Step 4, and twice in the Cowork section including "Sorry in advance but I'm gonna go all caps here"; the closing section is literally "Repeating one more time the core loop here for emphasis"). This is 'noticeably verbose; several unnecessary explanations or padded sections' rather than the mostly-efficient anchor at 3, though it does not explain concepts Claude already knows, keeping it above the anchor at 1.

2 / 5

Actionability

Guidance is largely executable: copy-paste subagent prompt templates ("Skill path: <path-to-skill> ... Save outputs to: <workspace>/iteration-<N>/eval-<ID>/with_skill/outputs/"), concrete JSON blocks for evals.json, eval_metadata.json, timing.json, and feedback.json, runnable commands ("python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>", "python -m scripts.run_loop --eval-set ... --max-iterations 5"), and exact field-name requirements ("must use the fields text, passed, and evidence"). Not a 5 because the central viewer command targets a path absent from the bundle ("<skill-creator-path>/eval-viewer/generate_review.py" — no eval-viewer/ directory exists, and scripts/generate_report.py is a different tool), so a key command fails as written; this fits 'mostly executable guidance; concrete code or commands with minor gaps'.

4 / 5

Workflow Clarity

The full loop is clearly sequenced (Capture Intent → Interview → Write SKILL.md → Test Cases → Steps 1-5 run/evaluate → Improve → iterate) with explicit validation checkpoints and feedback loops: user confirms test cases before running, assertions are graded per run, an analyst pass reads benchmark data, the user reviews in the viewer before improvements are applied, and "Keep going until: The user says they're happy / The feedback is all empty / You're not making meaningful progress" defines exit criteria. There is even a checklist instruction ("Please add steps to your TodoList... to make sure you don't forget") and error-recovery handling for alternate environments. Matches 'clear sequence with explicit validation steps; feedback loops for error recovery; checklists'.

5 / 5

Progressive Disclosure

Structure is decent — a dedicated "Reference files" section lists agents/grader.md, agents/comparator.md, agents/analyzer.md and references/schemas.md with one-line descriptions, and pointers are one level deep — but scoring against the actual bundle shows 4 of the referenced paths are dangling: there is no agents/ directory at all and no eval-viewer/generate_review.py, so following the pointers breaks. Only references/schemas.md and assets/eval_review.html exist. Missing referenced files plus Claude.ai/Cowork sections inlined at ~486 lines (near the skill's own 500-line guidance) fits 'some structure but could be better organized'; it is above the anchor at 2 (references are clearly signaled, not buried) and below the anchor at 4 ('minor organization gaps' understates four broken pointers).

3 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it answers both what the skill does and when to use it with explicit, natural trigger phrasing, and occupies a distinct authoring/eval niche. Small gains remain in adding a few natural synonyms and naming the concrete outputs of the workflow.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "Create new skills, modify and improve existing skills, and measure skill performance ... run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy" — naming what is done in each case. It falls short of the comprehensive anchor at 5 because it never mentions concrete artifacts the workflow produces (e.g., .skill packaging, test-case workspaces, benchmark reports), leaving minor coverage gaps.

4 / 5

Completeness

Both halves are explicit: the what is "Create new skills, modify and improve existing skills, and measure skill performance", and the when is an explicit clause enumerating trigger situations — "Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy." This mirrors the anchor example's what-plus-when structure with concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural user phrasings are well covered: "create a skill from scratch", "edit, or optimize an existing skill", "run evals", "benchmark skill performance", "optimize a skill's description". It misses a few natural synonyms and concrete nouns users would say — e.g., "SKILL.md", "skill authoring", "test a skill" appears only via "run evals to test a skill" — so it fits 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym-plus-extension coverage of the anchor at 5.

4 / 5

Distinctiveness Conflict Risk

The skill-authoring/eval niche is clear and distinct — "create a skill from scratch", "run evals to test a skill", "optimize a skill's description for better triggering accuracy" are phrases no general-purpose skill would claim. Minor overlap risk remains with generic evaluation/benchmarking or prompt-optimization skills, so it sits at 'mostly distinct; minor overlap risk' rather than the fully clear-cut niche of the anchor at 5.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
OpenBMB/PilotDeck
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.