CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-test

Validate skill files for structural compliance and behavioral correctness. Four modes: static linter, spec, category rubric, audit.

60

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-test/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, precisely validated multi-mode instruction set with explicit verdict semantics and error handling for every branch. Its weakness is structure at the file level: everything is inlined in one long SKILL.md rather than split across per-mode reference files, and a few passages are duplicated.

Suggestions

Split each mode's detail (static checks, spec evaluation rules, audit table logic, report templates) into per-mode reference files under references/ and keep SKILL.md as the routing overview, so an invocation of one mode does not carry the other three.

De-duplicate the NOT ASSESSED ranking explanation (stated in both Phase 2B Step 3 and Phase 2D Step 5) — state the ranking once and reference it.

Fix the phase numbering so document order matches labels (audit currently appears as Phase 2C after category's Phase 2D).

DimensionReasoningScore

Conciseness

The body is dense rule-setting with no padding explaining concepts Claude already knows — every section defines checks, verdicts, or edge cases. Minor over-explanation could be trimmed: the design-justification blockquote under Check 3 ("A linter that produces false failures… stops being trusted") and the NOT ASSESSED ranking being explained twice (Phase 2B Step 3 and Phase 2D Step 5).

4 / 5

Actionability

Fully executable instruction: exact glob patterns (`.claude/skills/*/SKILL.md`), literal per-check rules, copy-paste report templates with fixed verdict lines, exact error messages to print ("'…' is neither a skill…"), and concrete catalog fields (`last_spec:`, `last_category_result:`). Nothing is left as vague direction.

5 / 5

Workflow Clarity

Clear phased sequence with explicit validation everywhere: per-check PASS/WARN/FAIL verdicts, worst-case aggregation order (FAIL, PARTIAL, NOT ASSESSED, PASS), error-recovery messages for every missing-file case, ask-before-write gates ("May I write these results…"), and a denominator rule for the batch `static all` run. The one blemish — sections ordered 2A, 2B, 2D, 2C (audit/category labels swapped vs. document order) — does not break the sequence since each mode is independently entered.

5 / 5

Progressive Disclosure

Good internal section structure with clearly headed phases, but the entire four-mode specification — report templates, spec-evaluation rules, category rubric handling — is inlined in one ~530-line SKILL.md with no bundle files. Per-mode detail (e.g. the full spec-mode evaluation rules or the report templates) would fit naturally in separate reference files; as written, all content a model must load is loaded for every invocation.

3 / 5

Total

17

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, third-person, appropriately terse description that clearly names the domain and its four modes. Its main weakness is the complete absence of "when to use" trigger guidance, which both caps completeness and leaves trigger-term coverage thin.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks to test, lint, or validate a skill or agent, or wants a coverage audit of the skill catalog."

Include natural user phrasings and synonyms ("test my skills", "check skill compliance", "lint skills", "spec coverage") so the description matches what users actually say.

Briefly state each mode's purpose (e.g. "static: structural lint; spec: behavioral assertions; category: rubric metrics; audit: coverage report") to close the specificity gap.

DimensionReasoningScore

Specificity

Names the domain ("skill files") and several concrete capabilities — "structural compliance", "behavioral correctness", and four named modes ("static linter, spec, category rubric, audit"). It stops short of anchor 5 because it never states what each mode actually does (lint checks, spec assertion evaluation, rubric metrics, coverage report).

4 / 5

Completeness

Has a clear "what" (validate skills for structural compliance and behavioral correctness, in four modes) but "when" is entirely absent — no "Use when…" or equivalent trigger guidance, capping completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Some relevant keywords ("validate", "skill files", "linter", "audit") but missing the natural phrases a user would say — "test my skills", "check skill compliance", "lint", "skill coverage" — and no synonyms or variations. Not anchor 2 because the present keywords are domain-relevant rather than purely generic.

3 / 5

Distinctiveness Conflict Risk

Clear niche — validating a project's own skills and agents — with distinct terms ("static linter", "spec", "category rubric", "audit") that few other skills would claim. Minor overlap risk with adjacent QA/review skills keeps it below anchor 5.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (527 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
Donchitos/Claude-Code-Game-Studios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.