CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-benchmark-suite

Agent skill for benchmark-suite - invoke with $agent-benchmark-suite

52

2.17x
Quality

30%

Does it follow best practices?

Impact

89%

2.17x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/agent-benchmark-suite/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

32%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a monolithic catalog of non-executable class pseudocode with a duplicated frontmatter block, plus one genuinely useful section of CLI commands. It provides no sequenced workflow or validation checkpoints for what are destructive/batch operations, and it makes no use of progressive disclosure — everything, including material that belongs in reference files, is inlined. It should be reduced to a concise overview and operational procedure, with the class implementations moved to reference files or deleted.

Suggestions

Replace the pseudocode class catalogs with a concise step-by-step procedure anchored on the existing CLI commands (run benchmark -> compare baseline -> detect regression -> validate), including explicit validation checkpoints before declaring success, since these are long-running batch operations.

Move (or remove) the ~600 lines of class scaffolding and benchmark definitions into reference files (e.g. references/benchmarks.md) linked one level deep from a short overview, and delete the duplicated YAML frontmatter block at the top of the body.

Cut speculative and non-executable material — undefined classes and the mcp.benchmark_run/mcp.metrics_collect calls that have no real backing tooling — keeping only guidance the agent can actually execute.

DimensionReasoningScore

Conciseness

The ~650-line body is dominated by ~600 lines of illustrative class scaffolding (ComprehensiveBenchmarkSuite, RegressionDetector, AutomatedPerformanceTester, PerformanceValidator) encoding generic orchestration patterns Claude could generate itself, padded further by a duplicated YAML header block at the top of the body. It does not explain known concepts (keeping it above anchor 1) but is noticeably, heavily padded — anchor 2.

2 / 5

Actionability

The 'Operational Commands' section provides real, copy-pasteable CLI commands (e.g. 'npx claude-flow benchmark-run --suite comprehensive --duration 300', 'detect-regression --current <results> --historical <data>'), but the bulk of the body is pseudocode whose dependencies are undefined (ThroughputBenchmark, this.warmup, and a nonexistent mcp.benchmark_run/mcp.metrics_collect API). This matches anchor 3: 'Some concrete guidance but incomplete; pseudocode instead of executable code.'

3 / 5

Workflow Clarity

The body is a capability catalog, not a sequenced procedure — there is no ordered workflow the agent can follow, and validation exists only as descriptions inside pseudocode (a ResultValidator class, 'Validate results in real-time') rather than actionable checkpoints. Batch operations (5-minute benchmark suites, stress tests pushed to breaking points) lack any real validation loop, capping this at 3; the actual state matches anchor 2 ('steps poorly defined; validation absent').

2 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent) and the body references none; ~600 lines of class implementations and benchmark definitions that clearly belong in separate reference files are inlined in a single monolithic SKILL.md. Section headers exist, but with zero external references and everything inlined this fits anchor 2 ('content that clearly belongs in separate files is inlined') more than anchor 3, which assumes references are at least present.

2 / 5

Total

9

/

20

Passed

Description

28%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a meta placeholder ('Agent skill for benchmark-suite - invoke with $agent-benchmark-suite') rather than a capability statement: it names the domain but says nothing about what the skill does or when to use it. It fails both the 'what' and 'when' tests and offers almost no natural trigger terms. A rewrite stating concrete capabilities (performance benchmarking, regression detection, SLA validation) with an explicit 'Use when...' clause is needed.

Suggestions

State concrete capabilities in third person, e.g. 'Runs performance benchmark suites (throughput, latency, scalability), detects performance regressions against baselines, and validates results against SLA criteria' instead of the generic 'Agent skill for benchmark-suite'.

Add an explicit trigger clause: 'Use when the user asks to benchmark performance, compare against a baseline, detect performance regressions, or validate SLA/latency targets'.

Drop the meta invocation phrasing ('invoke with $agent-benchmark-suite') and include natural user vocabulary such as 'benchmark', 'performance test', 'load test', and 'regression detection' so the skill triggers on real requests.

DimensionReasoningScore

Specificity

The description only says 'Agent skill for benchmark-suite' — it names the domain but lists no concrete actions (nothing about benchmarking, regression detection, or validation). It matches anchor 2 ('Names the domain but actions are minimal or generic'); not 3 because no concrete action is actually stated, and not 1 because the domain is at least named.

2 / 5

Completeness

The 'what' is vague ('Agent skill for benchmark-suite') and the 'when' is entirely missing — there is no 'Use when...' clause or equivalent trigger guidance, which alone would cap this dimension at 3. It best matches anchor 2 ('Has a vague what and no when').

2 / 5

Trigger Term Quality

The only terms present are the technical identifiers 'benchmark-suite' and '$agent-benchmark-suite'; natural phrases a user would say ('benchmark', 'performance testing', 'regression', 'load test') are absent. This fits anchor 2 ('one or two generic keywords; missing the natural phrases users say').

2 / 5

Distinctiveness Conflict Risk

The benchmarking niche is somewhat distinct from unrelated skills, but the description would overlap with any performance-monitoring or optimization skill, and the meta 'invoke with $agent-benchmark-suite' phrasing does no discriminating work. Matches anchor 3 ('somewhat specific but could still overlap with similar skills').

3 / 5

Total

9

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (670 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
ruvnet/ruflo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.