CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-capability-analyzer

Runs the description-drift experiment — spawns all Claude Code agents simultaneously to collect self-reported capabilities, then compares them against static frontmatter descriptions to reveal how reliable orchestrator routing based on descriptions actually is. Use when measuring description drift across the agent fleet, re-running the capability collection experiment, analyzing a specific agent's self-reported capabilities, or auditing whether frontmatter descriptions accurately reflect agent behavior.

64

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/plugin-creator/skills/agent-capability-analyzer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with concrete commands and clear mode selection, but it carries substantial redundancy (the flowchart and prose describe the same modes twice), omits validation for a batch operation, and its most-critical referenced file (the capabilities prompt template) is missing from the bundle — with a setup section that also names the wrong dependency (level npm package vs the node:sqlite built-in the scripts actually use).

Suggestions

Remove the mermaid flowchart or the prose mode descriptions — they duplicate each other; keep one and cut the ~25 redundant lines to improve token efficiency.

Fix the bundle references: create resources/describe-your-capabilities.template.md (or correct its path), and delete the 'Setup (First Time Only)' section's `level` npm package instructions — the scripts use Node's built-in node:sqlite (Node >= 22.5.0), so no package install is needed.

Add validation checkpoints before the dump step: verify each spawned agent's entry exists in the store (e.g. re-run the lookup for each agent-id) and specify what to do when a Task fails or returns no tagged content, since the experiment batches writes from all agents simultaneously.

DimensionReasoningScore

Conciseness

The 25-line mermaid flowchart fully duplicates the prose that immediately follows it ('Single-agent mode — ...', 'Multi-agent mode — ...', 'Template usage — ...'), and the Full Experiment Workflow restates the invocation-mode steps again. This is more than the minor over-explanation of anchor 4, but the rest of the body is efficient, matching anchor 3.

3 / 5

Actionability

Concrete, executable commands throughout: 'node $CLAUDE_PLUGIN_ROOT/scripts/update-agent-map.mjs --name "agent-id" --capabilities …'', the dump command, the populate script, a node -e lookup one-liner, and a concrete report output template. Not anchor 5 because paths are inconsistent ('$CLAUDE_PLUGIN_ROOT/resources/…' vs './resources/…') and the pivotal template file reference does not resolve.

4 / 5

Workflow Clarity

The sequence (seed descriptions → read template → spawn agents → wait for all → dump → gap analysis) is clearly ordered across both modes, but this batch operation has no validation steps: no check that every spawned agent's entry landed in the store before the dump, and no handling of failed or non-responsive Tasks. Per the rubric's cap, a batch workflow without validation cannot score above 3.

3 / 5

Progressive Disclosure

Section structure is good and the Resources section lists bundle files one level deep, but scoring against the actual bundle reveals the core referenced artifact './resources/describe-your-capabilities.template.md' does not exist in the bundle (there is no resources/ directory) — a dangling reference to the exact prompt the whole workflow depends on. Better organized than anchor 2, but a broken primary reference keeps it below anchor 4.

3 / 5

Total

13

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person actions, an explicit 'Use when...' clause with multiple natural trigger scenarios, and a highly distinctive niche. The only weakness is slightly incomplete synonym coverage in its trigger terms.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'spawns all Claude Code agents simultaneously to collect self-reported capabilities', 'compares them against static frontmatter descriptions' — and covers the full workflow including single-agent analysis ('analyzing a specific agent's self-reported capabilities'). Anchor 4 requires minor coverage gaps; both invocation modes are represented, so anchor 5 fits best.

5 / 5

Completeness

It explicitly answers both questions: the 'what' ('spawns all Claude Code agents... then compares them against static frontmatter descriptions') and the 'when' with concrete trigger phrases ('Use when measuring description drift across the agent fleet, re-running the capability collection experiment, analyzing a specific agent's self-reported capabilities, or auditing...'). This matches anchor 5 exactly.

5 / 5

Trigger Term Quality

Good natural keyword coverage: 'description drift', 'agent fleet', 'self-reported capabilities', 'auditing whether frontmatter descriptions accurately reflect agent behavior'. It falls short of anchor 5 because common synonyms and variations (e.g. 'stale descriptions', 'agent routing', 'mismatched descriptions') are absent.

4 / 5

Distinctiveness Conflict Risk

A clear niche (description-drift auditing of an agent fleet) with distinct trigger terms; no realistic overlap with generic skills. Voice is third person ('Runs', 'spawns'), so no specificity penalty applies.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Jamie-BitFlight/claude_skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.