CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-performance-benchmarker

Agent skill for performance-benchmarker - invoke with $agent-performance-benchmarker

48

2.89x
Quality

26%

Does it follow best practices?

Impact

81%

2.89x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-performance-benchmarker/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

25%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a large non-executable JavaScript blueprint: concrete in algorithm design but dependent on phantom classes, with no runnable entry point, no workflow for the agent to follow, and no progressive disclosure into reference files. It reads as an auto-generated architecture sketch rather than operational skill guidance.

Suggestions

Move the implementation code into scripts/ files (or trim to a concise overview) and keep SKILL.md to purpose, entry-point invocation, and links to those files.

Define or stub the phantom dependencies (BenchmarkSuite, MetricsCollector, SystemMonitor, PerformanceModel, mcpTools) or reduce the code to a small runnable example against a real interface.

Replace the 'Core Responsibilities' list with a numbered execution workflow including validation checkpoints (e.g. verify environment, confirm suite registration, check results before applying optimizations).

DimensionReasoningScore

Conciseness

The body is a ~27KB monolithic dump of generic implementation scaffolding (console.log progress messages, comments like '// Initialize monitoring systems', speculative classes) that Claude could generate itself; nothing in it adds knowledge Claude lacks, matching anchor 1 ('severely verbose; heavily padded') rather than 2 which requires only 'several' padded sections.

1 / 5

Actionability

The code contains concrete algorithmic detail (adaptive rate ramp-up with success-rate thresholds, percentile/phase latency analysis, revert-if-improvement-under-5%) but every class depends on undefined infrastructure (TimeSeriesDatabase, MetricsCollector, SystemMonitor, PerformanceModel, this.mcpTools), making it detailed pseudocode rather than executable code — anchor 3, not 4 because nothing can actually run without the missing dependencies.

3 / 5

Workflow Clarity

The body offers only a 'Core Responsibilities' list and an implied order inside runComprehensiveBenchmarks; there is no step sequence for executing a benchmark (no setup, protocol instantiation, or environment instructions) and no validation checkpoints, matching anchor 2 ('rough sequence present but many gaps; validation absent') rather than 3 which requires clearly listed steps.

2 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent) and 800+ lines of per-system implementation that clearly belongs in separate script files are fully inlined in SKILL.md, matching anchor 2 ('content that clearly belongs in separate files is inlined') rather than 1 because section headers do provide some navigational structure.

2 / 5

Total

8

/

20

Passed

Description

28%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is pure invocation boilerplate: it restates the skill name and an invoke token without saying what the skill does or when to use it. It would rarely trigger correctly and gives Claude no capability information to distinguish it from other benchmarking skills.

Suggestions

Rewrite the description to state the skill's actual function, e.g. 'Measures throughput, latency, resource usage, and scalability of distributed consensus protocols and recommends parameter tuning.'

Add an explicit trigger clause naming user-facing keywords, e.g. 'Use when benchmarking consensus protocols, profiling throughput or tail latency, or comparing Raft/Byzantine performance.'

Drop the meta-boilerplate ('Agent skill for... - invoke with $...') in favor of third-person capability statements.

DimensionReasoningScore

Specificity

The description ('Agent skill for performance-benchmarker - invoke with $agent-performance-benchmarker') names the domain but states zero concrete actions, restating the skill name twice; it is not score 1 because the domain is clearly identified, and not score 3 because no capability like 'measures throughput and latency' is ever mentioned.

2 / 5

Completeness

The 'what' is only a vague restatement of the skill name (no mention of benchmarking, throughput, latency, or consensus protocols) and the 'when' is entirely absent with no 'Use when...' clause, matching anchor 2 ('has a vague what and no when') and well below the guideline's cap of 3 for missing trigger guidance.

2 / 5

Trigger Term Quality

The only keyword is the jargon compound 'performance-benchmarker' plus generic boilerplate ('Agent skill for', 'invoke with'); no natural phrases a user would say such as 'benchmark', 'throughput', or 'profile consensus' appear, matching 'one or two generic keywords; missing the natural phrases users say' rather than 1 since one domain keyword is present.

2 / 5

Distinctiveness Conflict Risk

The compound name 'performance-benchmarker' is somewhat niche, but because the description omits the actual domain (distributed consensus protocols) it would collide with any generic performance- or benchmarking-related skill, matching 'somewhat specific but could still overlap with similar skills' rather than 4's 'minor overlap risk'.

3 / 5

Total

9

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (856 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
ruvnet/ruflo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.