CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-judge

Evaluate Agent Skill design quality against official specifications and best practices. Use when reviewing, auditing, or improving SKILL.md files and skill packages. Provides multi-dimensional scoring and actionable improvement suggestions.

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality, genuinely expert evaluation framework with strong actionability and a clear workflow, undermined primarily by verbosity and a monolithic 749-line structure with no progressive disclosure. The conceptual 'what is a skill' padding and lack of external reference files are the main drag on an otherwise excellent meta-skill.

Suggestions

Trim the 'What is a Skill?', 'Tool vs Skill', and paradigm/training-vs-education sections plus decorative ASCII boxes — the base model already knows these concepts and they dilute the expert evaluation content (raises conciseness).

Move the detailed per-dimension anchors, the Common Failure Patterns catalog, and the full report template into references/ files (e.g. dimensions.md, failure-patterns.md, report-template.md) with one-level-deep links, keeping SKILL.md under ~300 lines as a routing overview (raises progressive_disclosure).

Add explicit 'MANDATORY - READ ENTIRE FILE' triggers and a 'Do NOT load' guidance block so the split-out reference files are loaded only when that dimension is being scored, rather than left unused.

DimensionReasoningScore

Conciseness

At 749 lines it carries substantial genuine expert framework (8 scored dimensions with anchors, red/green flags, failure-pattern taxonomy) but is padded with conceptual explanation the base model already knows — 'What is a Skill?', the training-vs-skills paradigm, 'Tool vs Skill' table, token-as-public-good lectures, and decorative ASCII boxes — so it is mostly efficient but could be tightened considerably.

2 / 3

Actionability

Provides concrete executable guidance: an explicit 5-step Evaluation Protocol, per-dimension scoring tables with numeric anchors and example evidence, a grade scale, a full report template, and a quick-reference checklist — copy-ready for an instruction-only skill.

3 / 3

Workflow Clarity

The Evaluation Protocol is a clearly sequenced multi-step process (knowledge-delta scan → structure analysis → score each dimension → calculate total/grade → generate report) with explicit checklists and a structured output template.

3 / 3

Progressive Disclosure

Well-organized into headed sections, but it is a 749-line monolithic SKILL.md with no references/ directory; content that could be split out (detailed dimension anchors, failure-pattern catalog, report template, full checklist) is all inline rather than one-level-deep referenced files.

2 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-constructed description that answers what, when, and supplies searchable trigger keywords in third-person voice. It would reliably activate when a user asks to review, audit, or improve a SKILL.md.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Evaluate Agent Skill design quality', 'reviewing, auditing, or improving', 'multi-dimensional scoring', 'actionable improvement suggestions' — rather than vague language; matches the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

Clearly answers both WHAT ('Evaluate Agent Skill design quality... Provides multi-dimensional scoring and actionable improvement suggestions') and WHEN via an explicit 'Use when reviewing, auditing, or improving...' clause.

3 / 3

Trigger Term Quality

Includes natural terms users would actually say ('reviewing, auditing, or improving SKILL.md files', 'skill packages') plus searchable domain terms ('SKILL.md files', 'skill packages'), giving good keyword coverage.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (skill quality review/auditing) with distinct triggers unlikely to fire for other skills; uses third-person voice throughout with no first/second-person phrasing.

3 / 3

Total

12

/

12

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (753 lines); consider splitting into references/ and linking

Warning

relative_links

Relative link issues: 1 missing

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

13

/

16

Passed

Repository
shareAI-lab/Kode-CLI
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.