CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-performance-analyzer

Agent skill for performance-analyzer - invoke with $agent-performance-analyzer

55

1.00x
Quality

31%

Does it follow best practices?

Impact

99%

1.00x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-performance-analyzer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

35%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized but hollow capability brochure: it describes what a performance analyzer would do in generic terms without a single command, code snippet, tool, or measurement technique that would let Claude actually perform the analysis. Its clearest asset is the three-phase workflow outline and the report template; its biggest liabilities are padding, the stray duplicate frontmatter block, and the complete absence of executable guidance.

Suggestions

Replace the generic detection/optimization bullet lists with concrete commands or code (e.g., timing wrappers, profiling invocations, how to compute parallelization ratio from a run log) so the skill instructs rather than describes.

Move the bottleneck-pattern catalog, KPI definitions, and report format into a references/ file (e.g. PATTERNS.md, REPORT-TEMPLATE.md) and keep SKILL.md as a lean overview with clearly signaled links.

Delete the stray second YAML frontmatter block at the top of the body and add validation checkpoints to the workflow (e.g., 'confirm the metric baseline is statistically meaningful before recommending a change; re-measure after applying to verify the improvement').

DimensionReasoningScore

Conciseness

The ~200-line body is noticeably padded with knowledge Claude already has — generic bullets like 'Resource Constraints: CPU, memory, or I/O limitations', 'Algorithm improvements', and 'Caching strategies' add no skill-specific information. It also opens with a stray duplicated YAML frontmatter block (name, capabilities, hooks) that is dead weight in the body. This matches anchor 2 ('Noticeably verbose; several unnecessary explanations or padded sections') rather than 1, since there is no extended tutorial-style prose explaining basics.

2 / 5

Actionability

There is no executable code, command, tool name, or metric-gathering method anywhere — the 'Analysis Workflow' phases are abstract numbered hints ('Gather execution metrics', 'Compare against baselines') and the 'Optimization Examples' assert results ('70% reduction to 3 minutes') without showing how they were achieved. Only the markdown report template is semi-concrete. This matches anchor 2 ('Minimal concrete guidance; high-level hints but missing the specific steps to execute') and falls short of 3, which would require at least pseudocode-level or partially complete instructions.

2 / 5

Workflow Clarity

A clear three-phase sequence exists (Data Collection → Analysis → Recommendation) with five enumerated steps each, matching anchor 3 ('Steps listed but validation gaps; sequence present but checkpoints missing'). It cannot score 4: there are no validation checkpoints or error-recovery loops (e.g., how to verify a suspected bottleneck or handle insufficient metrics), and the steps are generic rather than tied to concrete commands. The destructive/batch cap does not apply since the skill is read-only analysis.

3 / 5

Progressive Disclosure

The body has a reasonable section structure (capabilities, workflow, patterns, metrics, examples), but at ~200 lines everything lives in one file — the bottleneck-pattern catalog, KPI definitions, and report format clearly belong in separate reference files — and no references exist at all. This matches anchor 3 ('Some structure but could be better organized; content that should be separate is inline'). The 5-scoring exception for under-50-line skills does not apply.

3 / 5

Total

10

/

20

Passed

Description

28%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a placeholder-quality label rather than a functional description: it tells the model how to invoke the skill but neither what it does nor when to use it. All four dimensions score at or near the bottom of the scale, with completeness being the most serious failure.

Suggestions

State concrete capabilities in third person, e.g. 'Identifies performance bottlenecks in workflows and agent coordination, profiles resource usage, and recommends parallelization or caching strategies.'

Add an explicit 'Use when...' clause with natural trigger phrases such as 'Use when tasks run slowly, execution time regresses, or the user mentions bottlenecks, latency, or profiling.'

Drop the '$agent-performance-analyzer' invocation instruction from the description — it consumes the description budget without contributing capability or trigger information.

DimensionReasoningScore

Specificity

The description 'Agent skill for performance-analyzer' names the domain but contains no action verbs whatsoever — 'invoke with $agent-performance-analyzer' is an invocation instruction, not a capability. It sits above anchor 1 ('Helps with documents') only because it names a domain, but below anchor 3 since zero concrete actions (analysis, bottleneck identification, etc.) are stated.

2 / 5

Completeness

The 'what' is only vaguely implied (a skill 'for performance-analyzer' says nothing about what it does), and the 'when' is entirely absent with no 'Use when...' clause or equivalent — a textbook match for anchor 2 ('Has a vague what and no when'). Per the judging guidelines, a missing 'Use when' clause would also cap completeness at 3 even if the what were strong, and here the what is weak as well.

2 / 5

Trigger Term Quality

The only keywords are 'performance-analyzer' and the '$agent-performance-analyzer' token — technical identifiers rather than natural phrases. Natural trigger terms users would say ('slow', 'bottleneck', 'latency', 'optimize workflow', 'profiling') are entirely missing, matching anchor 2's 'one or two generic keywords; missing the natural phrases users say'.

2 / 5

Distinctiveness Conflict Risk

Naming 'performance-analyzer' gives it a somewhat specific domain that would not collide with unrelated skills, matching anchor 3 ('Somewhat specific but could still overlap with similar skills'). It cannot score 4 because no concrete triggers or capability details are given, so it could easily fire for any performance, profiling, or optimization request.

3 / 5

Total

9

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ruvnet/ruflo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.