CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-creator

Create, refine, and benchmark agent skills. Use when building a new skill, updating an existing one, running evals, checking trigger quality, or improving a skill description.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-creator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-sequenced process skill with explicit validation checkpoints and iteration feedback loops, undermined by broken bundle references (agents/*.md and eval-viewer/generate_review.py are missing) and chatty token-padding that pushes the body past its own 500-line guidance. The workflow itself is exemplary.

Suggestions

Fix the broken bundle references: create the agents/ directory (grader.md, comparator.md, analyzer.md) and eval-viewer/generate_review.py, or repoint those instructions at the files that actually exist (e.g., scripts/generate_report.py) so the copy-paste commands run as written.

Trim the chatty padding — 'Cool? Cool.', 'Good luck!', the 'Communicating with the user' section, and the end-of-file verbatim repeat of the core loop — to reclaim tokens without losing any guidance.

Move the detailed Description Optimization section (trigger-eval query writing guidance and the full run_loop CLI flag list) into references/description-optimization.md with a short pointer, bringing the body comfortably under the 500-line limit it itself recommends.

DimensionReasoningScore

Conciseness

The bulk is dense, actionable process guidance, but padded with chatty filler ("Cool? Cool.", "Good luck!", "maybe literally, maybe even more who knows"), a "Communicating with the user" section of marginal value, and a verbatim repeat of the core loop ("Repeating one more time the core loop here for emphasis"). Matches anchor 3 (mostly efficient, some unnecessary explanation that could be tightened); not 2 because padding is stylistic rather than whole redundant sections, not 4 because several passages clearly fail to earn their tokens.

3 / 5

Actionability

Highly concrete throughout: exact JSON blocks (eval_metadata.json, timing.json, feedback.json), copy-paste bash commands ("python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>", "python -m scripts.run_loop --eval-set ..."), and ready-to-use subagent prompt templates. Held below 5 because several exact commands point at files absent from the bundle ("eval-viewer/generate_review.py", "agents/grader.md"), so they don't run as-is; well above 3 since nothing is pseudocode.

4 / 5

Workflow Clarity

The multi-step process is explicitly sequenced with validation checkpoints and feedback loops: "Wait to write test prompts until you've got this part ironed out", user sign-off on test cases and eval queries, "When the user tells you they're done, read feedback.json", the improve-then-rerun iteration loop ending only when "the feedback is all empty", plus error-recovery paths (headless "--static" fallback, kill $VIEWER_PID). Matches anchor 5; not 4 because checkpoints are explicit, not implicit.

5 / 5

Progressive Disclosure

Structure and signaling are good — references are one level deep with when-to-read guidance ("Read agents/grader.md ... when you need to spawn the relevant subagent") and references/schemas.md, assets/eval_review.html, and the scripts/ files all resolve. But scored against the actual bundle, the referenced agents/ directory (grader.md, comparator.md, analyzer.md) and eval-viewer/generate_review.py do not exist, and the ~507-line body inlines the entire detailed Description Optimization section (~80 lines of CLI flags and query-writing guidance) that belongs in a reference file. Matches anchor 3; not 4 because broken referenced paths and inlineable sections exceed 'minor organization gaps', not 2 because most paths resolve and navigation is clear.

3 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concise, with a clear what and an explicit, well-enumerated 'Use when' trigger clause covering build, update, eval, and description-improvement contexts. The only weakness is modest trigger-term synonym coverage and slight genericness of 'running evals'.

DimensionReasoningScore

Specificity

"Create, refine, and benchmark agent skills" names the domain plus three concrete actions, matching the 'several specific actions; minor gaps' anchor. It falls short of 5 because coverage has gaps (packaging, blind comparison, and description optimization are only implied via 'refine'); it exceeds 3 because more than 1-2 concrete actions are given.

4 / 5

Completeness

Explicitly answers both: the 'what' ("Create, refine, and benchmark agent skills") and a fully enumerated 'when' introduced by "Use when...". Matches the anchor 5 pattern exactly; not 4 because the 'when' clause is explicit and concrete rather than merely present.

5 / 5

Trigger Term Quality

"building a new skill, updating an existing one, running evals, checking trigger quality, or improving a skill description" gives good natural-phrase coverage a user would actually say. Not 5 because common synonyms and artifacts are missing (e.g., 'make a skill', 'write a skill', 'SKILL.md', 'package a skill'); not 3 because coverage is broad and varied rather than partial.

4 / 5

Distinctiveness Conflict Risk

"agent skills" plus "checking trigger quality" and "improving a skill description" form a clear niche, but generic terms like "running evals" and "benchmark" carry minor overlap risk with general model-eval or testing skills. Mostly distinct with minor overlap, matching anchor 4; not 5 because of that overlap, not 3 because the core triggers are unmistakably about skill creation.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (508 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
feiskyer/claude-code-settings
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.