CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-skill

Evaluate a local Codex skill in engineer-friendly terms. Use when the user says "evaluate this skill", "give me an analysis of the game dev skill", "audit this skill", "why did this score that way", "what should I fix first", or asks for a skill-specific report before benchmarking it.

65

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/plugin-eval/skills/evaluate-skill/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is concise and highly actionable with concrete, copy-paste-ready commands and a clear routed workflow. Its chief weakness is progressive disclosure: the one reference link it makes is broken, and several sections duplicate content already present in the description or workflow.

Suggestions

Fix or remove the broken reference '../../references/chat-first-workflows.md' (the file does not exist), or replace it with a real one-level-deep reference file.

Drop the standalone 'Commands' block or the 'Chat Requests To Recognize' list since both duplicate the inline workflow commands and the description triggers, trimming tokens.

Add a light validation step to the workflow (e.g. confirm plugin-eval is installed / the skill path resolves before running analyze) to add a feedback checkpoint.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence with no concept over-explanation, but the 'Commands' block repeats commands already shown inline in the workflow and 'Chat Requests To Recognize' duplicates triggers already in the description, giving minor trim opportunities; not a 5 because of that duplication, and not a 3 because the prose is genuinely efficient rather than padded.

4 / 5

Actionability

It provides copy-paste-ready, fully executable commands with concrete flags (e.g. 'plugin-eval analyze <skill-path> --format markdown') across the common cases (start, analyze, explain-budget, measurement-plan, init-benchmark, benchmark --dry-run), matching the anchor for fully executable guidance that covers common cases.

5 / 5

Workflow Clarity

A clearly sequenced 10-step workflow with explicit conditional routing ('If the user says ...', 'If the user wants ...') and unambiguous steps; not a 5 because there are no explicit validation/feedback checkpoints, and not a 3 because the sequence is coherent and well-defined with no real gaps (the destructive-operation cap does not apply since evaluation/benchmarking is non-destructive).

4 / 5

Progressive Disclosure

The body is well-sectioned (Workflow, Priorities, Commands, Reference) and the single reference is clearly signaled, but the only external reference '../../references/chat-first-workflows.md' points to a file that does not exist, breaking navigation to the deeper content; not a 4 because a broken reference is more than a minor organization gap, and not a 2 because the in-body structure itself is reasonable rather than minimal.

3 / 5

Total

16

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that clearly answers both what it does and when to use it, with a focused niche and several natural trigger phrases. Its main limitation is that it names only one concrete action, leaving capability coverage a touch thin.

Suggestions

Add one or two more concrete actions (e.g. 'score, benchmark, and recommend fixes for a local Codex skill') to lift specificity from a single action to several.

Include a compact shorthand trigger such as 'score this skill' to round out keyword coverage toward comprehensive.

DimensionReasoningScore

Specificity

It names the domain ('a local Codex skill') and one concrete action ('Evaluate ... in engineer-friendly terms') but offers no second distinct action, matching the anchor for 1-2 concrete actions without comprehensive coverage; not a 4 because there aren't 'several specific actions,' and not a 2 because it goes beyond merely naming the domain.

3 / 5

Completeness

It explicitly states what ('Evaluate a local Codex skill in engineer-friendly terms') and when ('Use when the user says ...') with concrete trigger phrases, matching the anchor that clearly and explicitly answers both what and when.

5 / 5

Trigger Term Quality

It lists six natural phrasings a user would actually say ('evaluate this skill', 'audit this skill', 'why did this skill score that way', 'what should I fix first') giving good keyword coverage with synonyms; not a 5 because coverage is not exhaustive (no shorthand like 'score this skill') and not a 3 because the terms go well beyond a single generic keyword.

4 / 5

Distinctiveness Conflict Risk

The Codex-skill-evaluation niche plus scoring/benchmarking-specific triggers make it mostly distinct with only minor overlap risk against generic review/audit skills; not a 5 because phrases like 'evaluate this skill' and 'audit this skill' could plausibly match other review skills, and not a 3 because the domain is clearly scoped rather than broad.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.