CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, update or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./SKILLs/skill-creator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body excels at workflow clarity and actionable detail — concrete commands, JSON schemas, checkpointed steps, and environment-specific adaptations — but is undermined by chatty padding that inflates token cost and by broken references to `agents/*.md` and `eval-viewer/generate_review.py` that don't exist in the bundled files. Trimming the conversational asides and reconciling paths with the actual bundle would lift both weak dimensions.

Suggestions

Cut the conversational padding — "Cool? Cool.", the plumber/grandparent anecdote, the "billions a year in economic value" aside, and the all-caps apology — and state the core loop once instead of three times (intro, iteration-loop section, closing recap).

Reconcile referenced paths with the actual bundle: `agents/grader.md`, `agents/comparator.md`, and `agents/analyzer.md` are missing entirely, and `eval-viewer/generate_review.py` should be `scripts/generate_report.py` (or the files should be added).

Move the Claude.ai- and Cowork-specific adaptation sections into a short reference file linked from a one-line pointer, reducing the main body's length while keeping the environment guidance available on demand.

DimensionReasoningScore

Conciseness

Several padded sections: "Cool? Cool.", the plumber/grandparent anecdote about who uses terminals, "we are trying to create billions a year in economic value here!", "Sorry in advance but I'm gonna go all caps here", and the core loop restated three times (intro bullets, iteration-loop section, and the closing recap). This is noticeably verbose with multiple unnecessary sections, though it avoids explaining concepts Claude already knows, so it does not reach the 'severely verbose' bottom anchor.

2 / 5

Actionability

Guidance is mostly executable: copy-paste-ready subagent prompt templates, exact JSON formats (eval_metadata.json, timing.json, feedback.json), and concrete bash commands (`python -m scripts.aggregate_benchmark`, the nohup viewer invocation, `python -m scripts.run_loop`). It falls short of fully executable because key referenced paths dangle: `agents/grader.md`, `agents/analyzer.md`, `agents/comparator.md`, and `eval-viewer/generate_review.py` are invoked but not present in the bundle (the viewer script actually lives at `scripts/generate_report.py`).

4 / 5

Workflow Clarity

The multi-step process is clearly sequenced with explicit validation checkpoints: numbered Steps 1-5 for run/assertion/timing/grading/feedback, grading against assertions, benchmark aggregation, an analyst pass, a user-review gate, and explicit stop conditions ("user says they are happy / feedback is all empty / not making meaningful progress"). Batch operations do have validation, so no cap applies; error recovery via feedback loops is present throughout.

5 / 5

Progressive Disclosure

References are clearly signaled and there is a dedicated Reference-files section, and `references/schemas.md`, `assets/eval_review.html`, and the `scripts/` modules all resolve. However, scored against the actual bundle, four referenced paths (the `agents/` directory and `eval-viewer/generate_review.py`) do not exist, breaking navigation for the central grading and viewer steps — more than the 'minor organization gaps' of a 4.

3 / 5

Total

14

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does and when to use it, with a pushy, explicit trigger clause and lifecycle-wide action coverage. Weaknesses are minor: a few natural synonym phrasings and the .skill extension are missing, and eval/benchmark language carries slight overlap risk with general testing skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions spanning the full skill lifecycle: "Create new skills, modify and improve existing skills, and measure skill performance" plus "run evals", "benchmark skill performance with variance analysis", and "optimize a skill's description for better triggering accuracy". This is comprehensive coverage of the domain, matching the top anchor rather than the 'minor gaps' of a 4.

5 / 5

Completeness

It explicitly answers both questions: the first sentence states what the skill does, and "Use when users want to create a skill from scratch, update or optimize an existing skill, run evals to test a skill, benchmark..." provides concrete, explicit trigger phrases. This matches the top anchor exactly.

5 / 5

Trigger Term Quality

Trigger phrases like "create a skill from scratch", "update or optimize an existing skill", "run evals", and "benchmark skill performance" are natural user phrasings. It falls short of a 5 because common synonyms ("make a skill", "write a skill") and the ".skill" file extension referenced by the workflow are missing.

4 / 5

Distinctiveness Conflict Risk

Every clause is anchored on "skill", giving a clear niche with distinct triggers. Minor overlap risk remains: "run evals" and "benchmark skill performance" could pull in general testing/eval skills, so it fits 'mostly distinct; minor overlap risk' rather than the minimal-conflict top anchor.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
netease-youdao/LobsterAI
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.