CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-cleaner

Codex/OpenClaw skill audit: live budget, usage, duplicates, compact descriptions.

59

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-cleaner/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, lean, highly actionable skill body: executable commands with useful variants, a clear report-reading order, and explicit safety rules around destructive cleanup. The only real improvement is structural — moving the script-internals detail in 'Analyzer Notes' into a reference file and adding a post-cleanup re-validation step to the workflow.

DimensionReasoningScore

Conciseness

The body is dense and lean — no concept explanations Claude already knows, no padded prose; every section carries operational information (commands, report-reading order, decision rules). Score 4 rather than 5 because 'Analyzer Notes' packs in implementation details of the script (e.g. 'token cost `ceil(utf8_bytes / 4)`, then full descriptions -> equal description truncation -> omitted minimum lines', realpath-dedup internals) that could be trimmed or moved out of the always-loaded body. Clearly above anchor 3 ('some unnecessary explanation') since nothing here is filler.

4 / 5

Actionability

The guidance is fully executable: a copy-paste-ready base command, six concrete flag variants covering common cases (--no-logs, --months 6 --deep-logs, --context-tokens 272000 --budget-percent 2, custom --root), an explicit report-reading order keyed to named report sections, and concrete keep/delete decision rules. This matches anchor 5 ('fully executable; copy-paste ready code or commands; specific examples cover the common cases'); anchor 4 would imply minor gaps in coverage, and none are evident.

5 / 5

Workflow Clarity

The three-step workflow (run analyzer -> read report sections in order -> pre-deletion verification) is clearly sequenced, and this batch/destructive skill does include validation checkpoints: 'Verify the kept copy exists and is loaded', 'Suggest first; edit only when the user asks', and the guard against deleting ignored/untracked dirs without confirmation. Score 4 rather than 5 because there is no post-action feedback loop (e.g. re-run the analyzer after cleanup to confirm budget shrank) — the checkpoints are pre-flight checks rather than a validate/re-validate cycle, leaving a minor validation gap per anchor 4.

4 / 5

Progressive Disclosure

Structure is good: the body is an overview that delegates all code to a real one-level-deep bundle file (scripts/skill-cleaner.ts and its test exist in the bundle), with well-signaled sections (Workflow, Analyzer Notes, Output Policy). Score 4 rather than 5 because the 'Analyzer Notes' section (~13 dense lines of script internals) is content that could live in a reference file, keeping the always-loaded body tighter — a minor organization gap per anchor 4; it does not fall to anchor 3 since what is inline is operational rather than bulk reference material.

4 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is admirably concise and domain-specific, but it reads like a terse label rather than a trigger description: it enumerates output areas as noun fragments without stating actions, and it omits any 'Use when...' guidance. It would benefit most from an explicit trigger clause and verb-based capability statements with common synonyms (unused, remove, disable, clean up).

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user wants to trim skill prompt budget, find duplicate or unused skills, or decide which Codex/OpenClaw skills or plugins to remove.'

Convert noun fragments into concrete third-person actions, e.g. 'Audits live Codex skill inventory against the 2% prompt budget, reports usage from session logs, flags duplicate and unused skills, and suggests compacted descriptions.'

Include common synonyms users would naturally say — 'unused', 'remove', 'disable', 'clean up' — alongside 'duplicates' and 'budget' to broaden trigger coverage.

DimensionReasoningScore

Specificity

The description names the domain ("Codex/OpenClaw skill audit") and several specific output areas ("live budget, usage, duplicates, compact descriptions"), but these are terse noun fragments rather than concrete actions — nothing says what the skill actually *does* (analyze, report, suggest deletions). This matches anchor 3 ('names domain and 1-2 concrete actions, but not comprehensive') better than anchor 4, which expects a list of specific actions with only minor gaps; it does not fall to anchor 2 because the enumerated items are domain-specific, not generic.

3 / 5

Completeness

The 'what' is present but compressed into a fragment list (audits skills for budget, usage, duplicates, compact descriptions), and the 'when' is entirely absent — there is no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Not score 2 because the 'what' is reasonably clear rather than vague; not score 4 because no 'when' exists at all.

3 / 5

Trigger Term Quality

Relevant keywords are present ("skill audit", "budget", "duplicates", "usage", "compact descriptions", "Codex/OpenClaw"), but common natural phrasings users would actually say are missing: "clean up skills", "unused skills", "remove/disable skills", "trim". Anchor 3 ('some relevant keywords but missing common variations or synonyms') is the closest fit; anchor 4 ('good keyword coverage; a few natural terms missing') would require broader coverage of everyday synonyms, and anchor 2 would imply generic-only keywords, which is not the case.

3 / 5

Distinctiveness Conflict Risk

The niche is clear and fairly distinct — auditing Codex/OpenClaw skill inventories for budget pressure, duplicates, and usage — with domain-specific terms ("live budget", "duplicates") that would rarely trigger the wrong skill. Minor overlap risk remains with generic cleanup/organization skills and with code-review/dedup skills. Anchor 4 ('mostly distinct; minor overlap risk') fits best; anchor 5 would require trigger phrasing that fully separates it from adjacent skills.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
steipete/agent-scripts
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.