CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-doctor

Grades agent skills by scoring agent conversations for efficiency, code quality, procedure compliance, and verbosity, then drafts concrete skill edits and a shareable report. Use when the user wants their agent setup graded from real conversation history, or asks which of their installed skills are actually working.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tightly written, well-sequenced operational skill with excellent conciseness and genuine validation gates for a batch workflow. Its main defects are executable-level: the collector command uses placeholder arguments, and the four scorer rubric files referenced in Step 2 are missing from the bundle, leaving the central scoring step unresolvable as shipped.

Suggestions

Add the missing scorers/ bundle files (efficiency.md, code-quality.md, procedure-compliance.md, verbosity.md) referenced in Step 2, or inline their label tables — as bundled, all four '$SKILL_ROOT/scorers/*.md' paths are dead and the scoring phase cannot proceed.

Replace the '<conversation-scope arguments>' and '<skill-scope arguments>' placeholders in the collect_sessions.py invocation with the concrete flag forms (e.g. '--repo "$REPO"', '--all-conversations', '[--include-global-skills]') so the command is copy-paste executable.

Provide a short fallback in Step 2 for the missing-rubrics case (e.g. 'if a scorer file is absent, halt and tell the user') so the run degrades gracefully instead of silently improvising scoring criteria.

DimensionReasoningScore

Conciseness

The body is dense procedural instruction with essentially no padding: no concept explanations, no library surveys, no motivational text. Every section (harness gate, scope questions, collector flags, aggregation formulas, diff drafting, report template, output summary) carries operational content Claude does not already know. It matches anchor 5 ('lean and efficient; every token earns its place') rather than 4, which requires identifiable over-explanation — I could not find a passage that is trimmable without losing a real constraint.

5 / 5

Actionability

Most commands are copy-paste ready: 'git rev-parse --show-toplevel', the 'mktemp -d' scratch-dir creation, the render invocation, and 'diff -u <current> <proposed>'. The collector invocation, however, is a template with '<conversation-scope arguments>' placeholders (resolvable via the flag mapping listed immediately above, but not literally executable), and Step 2 tells the agent to 'Pass the following rubrics as context' — four '$SKILL_ROOT/scorers/*.md' paths that do not exist in the bundle, so the central scoring step cannot run as shipped. That is more than the 'minor gaps' of anchor 5 but well above anchor 3's pseudocode, since the real commands that are given are executable.

4 / 5

Workflow Clarity

Steps 0–6 are clearly sequenced with explicit validation checkpoints and error-recovery feedback loops for a batch operation: the harness gate before any history is read, 'If sessions_sampled is 0... and stop' with a concrete remedy (raise --days), 'If skills_found is 0, continue' with defined behavior, path validation before proceeding ('Expand and validate every path as a git repository'), and insufficient_evidence exclusion rules for the batch scoring. This matches anchor 5 ('clear sequence with explicit validation steps; feedback loops for error recovery') and avoids the batch-operation cap at 3 because validation is present throughout.

5 / 5

Progressive Disclosure

The SKILL.md is a well-organized overview that defers detail to one-level-deep, clearly signaled references: '$SKILL_ROOT/references/supported-harnesses.md' for harness specifics (twice, for the gate and the collector flags) and 'references/skill-improvements.md' for edit drafting, plus the scripts directory — all of which exist in the bundle. However, the four '$SKILL_ROOT/scorers/*.md' paths referenced in Step 2 do not exist ('scorers/' is absent from the bundle), so the skill's core rubric content is unreachable as shipped. This sits between anchor 4 ('references mostly clear; minor organization gaps') and anchor 3; because most referenced paths resolve and the structure itself is good, 4 fits better than 3, whose defining traits (unsignaled references, inline content that belongs in separate files) are absent.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description in third person that states a concrete multi-step capability set and pairs it with an explicit, naturally worded 'Use when' clause. The only gap is modest synonym coverage in the trigger phrasing. Its voice is consistently third person ('Grades', 'drafts'), satisfying the voice guideline.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions covering the whole pipeline: 'scoring agent conversations for efficiency, code quality, procedure compliance, and verbosity, then drafts concrete skill edits and a shareable report.' All four scoring dimensions plus both output artifacts are named explicitly. Anchor 5 is the best fit; 4 ('minor gaps in coverage') would understate it since no pipeline stage is omitted.

5 / 5

Completeness

It explicitly answers both questions: the 'what' is the grading/editing/report pipeline, and the 'when' is a literal 'Use when...' clause with two concrete trigger scenarios. This matches anchor 5 ('clearly and explicitly answers both what AND when with concrete trigger phrases') and is strictly stronger than anchor 4's 'when could be more explicit or specific'.

5 / 5

Trigger Term Quality

The 'Use when' clause contains natural user phrasings — 'wants their agent setup graded from real conversation history', 'asks which of their installed skills are actually working' — that a user would plausibly say verbatim. A few natural synonyms are missing (e.g. 'audit my skills', 'skill health', 'grade my skills'), so it matches anchor 4 ('good keyword coverage; a few natural terms missing') rather than 5's comprehensive synonym/extension coverage.

4 / 5

Distinctiveness Conflict Risk

'Grades agent skills by scoring agent conversations' plus 'installed skills are actually working' carves out a clear niche (auditing one's own agent/skill setup from local history) with triggers unlikely to fire for unrelated skills. It sits well below anchor-4 overlap risk because no closely competing skill category (e.g. code review) matches its trigger phrasing.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
warpdotdev/common-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.