CtrlK
BlogDocsLog inGet started
Tessl Logo

system-reviewer

Review the skill development ecosystem itself - assess ecosystem health, identify systemic issues, evaluate toolkit effectiveness, and recommend system-level improvements. Task-based operations for ecosystem assessment, toolkit evaluation, process review, and system optimization. Use when evaluating ecosystem health, identifying systemic improvements, optimizing the toolkit itself, or conducting meta-level ecosystem reviews.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/system-reviewer/SKILL.md

The canonical home for this skill is system-reviewer in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized and clearly delineates five operations with consistent structure, but it reads more like a framework outline than an executable skill: it tells Claude what to assess without how to measure anything, has no validation checkpoints, and inlines all detail (including a full example report and stale hardcoded counts) in a single long file. Organization is the strength; actionability and offloading are the gaps.

Suggestions

Add concrete data-gathering methods to each operation: which files, logs, or metrics sources to read (e.g., where build-time and rework data comes from, or a command/script to inventory skills per layer), so 'Measure Efficiency Gains' and 'Measure Process Metrics' become executable rather than aspirational.

Move the per-operation details and the ~50-line example health report into one-level-deep reference files (e.g., references/ecosystem-health-report.md), keeping SKILL.md as an overview with clearly signaled links.

Remove hardcoded stale counts ('39 skills originally', '23 skills currently', the dated 2025-11-07 example) from the process steps — instruct Claude to derive current counts from the actual ecosystem instead — and state the system-vs-individual distinction once instead of three times.

DimensionReasoningScore

Conciseness

The body is mostly operational rather than padded with concepts Claude already knows, but it could be tightened: the system-vs-individual distinction is stated three times ('Key Distinction' in Overview, Best Practice #3, and the Quick Reference section), the 5 operations are enumerated twice (Overview list and Quick Reference table), and Operation 4 hardcodes stale time-sensitive counts ('39 skills originally', '23 skills currently') directly in the process steps. It is not 'noticeably verbose with several padded sections' (anchor 2) since most content is directive, but the duplication and the ~50-line example report keep it below anchor 4.

3 / 5

Actionability

Each operation provides concrete checklists of what to examine ('Build time progression', 'Rework rate', 'Which tools used most?') and Operation 1 includes a full example output report that serves as a concrete template. However, the guidance stops short of explaining how to obtain the data — no commands, scripts, metric sources, or methods are given for measuring 'efficiency vs baseline' or 'blocker frequency' — so execution details are missing, matching anchor 3 ('some concrete guidance but incomplete').

3 / 5

Workflow Clarity

Every operation has a clearly numbered 5-step sequence with defined outputs and time estimates, which is good sequencing. But checkpoints are missing or implicit: no step verifies findings (e.g., how to validate a health score or cross-check gap analysis against the actual skill inventory), so it matches anchor 3 ('steps listed but validation gaps'). It is above anchor 2 because sequences are well defined with explicit outputs, and the destructive/batch cap does not apply since these are read-only assessment operations.

3 / 5

Progressive Disclosure

Section structure is clear (Overview, When to Use, Operations, Best Practices, Quick Reference), but the entire ~385-line body lives in one file with no bundle files at all — the detailed per-operation processes (each 40-60 lines) and the 50-line example health report are content that could live in one-level-deep reference files. This matches anchor 3 ('some structure but content that should be separate is inline'), and the under-50-lines exception explicitly does not apply to a file this long.

3 / 5

Total

12

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete meta-level capabilities and includes an explicit, trigger-phrase-bearing 'Use when' clause, so a user could reliably select this skill over siblings. The main weakness is the redundant middle sentence that re-lists the same capabilities as noun phrases without adding information or new trigger terms.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'assess ecosystem health, identify systemic issues, evaluate toolkit effectiveness, and recommend system-level improvements' — which cover the domain well. It falls short of a 5 because the middle sentence ('Task-based operations for ecosystem assessment, toolkit evaluation, process review, and system optimization') restates the same capabilities as redundant filler rather than adding new concrete actions.

4 / 5

Completeness

It explicitly answers both questions: the 'what' ('assess ecosystem health, identify systemic issues, evaluate toolkit effectiveness, and recommend system-level improvements') and a concrete 'when' ('Use when evaluating ecosystem health, identifying systemic improvements, optimizing the toolkit itself, or conducting meta-level ecosystem reviews'). This matches the anchor-5 pattern of explicit what-and-when with concrete trigger phrases.

5 / 5

Trigger Term Quality

Good natural keyword coverage: 'ecosystem health', 'systemic issues', 'toolkit evaluation', 'process review', 'system optimization', plus an explicit 'Use when' clause with gerund triggers. It misses anchor 5 because phrases like 'meta-level ecosystem reviews' are jargon users would rarely say, and common synonyms (e.g., 'skill system', 'review the toolkit') are absent.

4 / 5

Distinctiveness Conflict Risk

The niche is clear — 'the skill development ecosystem itself' and 'systemic issues' distinguish it from individual-skill reviewers — giving minimal conflict with unrelated skills. It does not reach 5 because the description still risks overlap with closely related review skills (e.g., a generic skill-reviewer), as 'system-level improvements' could plausibly trigger for individual skill review requests.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.