CtrlK
BlogDocsLog inGet started
Tessl Logo

http-load-profiler

Run stepped HTTP load tests with ab/wrk, ramping concurrency levels to collect p50/p90/p99 latency, detect performance inflection points, and recommend optimal concurrency. Triggered by requests like 'load test this URL', 'benchmark my API', 'find the max concurrency', or mentions of p99 latency, throughput saturation, or capacity planning.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Medium

Suggest reviewing before use

SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured tool skill with executable commands and accurate documentation of the bundled script. It loses points for duplicating the inflection-point rules across two sections and for lacking any validation/smoke-test checkpoint before ramping real traffic against a target.

Suggestions

Consolidate the four inflection-point rules: state them once (in 'Inflection Point Detection Algorithm') and have the Features bullet simply link or summarize, cutting the duplicated thresholds list.

Add a pre-ramp validation step, e.g. 'Verify the endpoint first: run one step at concurrency 1 and confirm a 200/valid response before continuing to higher steps' — this both satisfies the batch-operation safety concern and gives the workflow a checkpoint.

Consider moving the full JSON output example (or the human-readable sample) into a references/ file and keeping only a schema sketch inline, which would also trim the SKILL.md token budget.

DimensionReasoningScore

Conciseness

The body is mostly efficient — a lean parameters table, tight quick-start, terse notes — but the Features section restates all four inflection rules almost verbatim ('p99 latency accelerating… Throughput efficiency dropping… Throughput saturated… Error rate surging') and the 'Inflection Point Detection Algorithm' section repeats them a second time, plus the intro sentence duplicates the description. Not 2 because there is no tutorial-style padding or explanation of concepts Claude already knows; not 4 because the internal duplication is a real tightening opportunity, not a minor trim.

3 / 5

Actionability

Quick Start gives copy-paste-ready commands covering basic use, custom steps, engine override, request count, and JSON output; the Parameters table documents every flag with defaults; Prerequisites includes per-platform install commands. Common cases are fully covered with executable commands, matching the top anchor; it is not 4 because nothing is pseudocode or missing.

5 / 5

Workflow Clarity

The single-command flow is unambiguous, but load testing is a batch operation that generates real traffic against a service (the doc itself warns 'do not run against production services without authorization') and the workflow has no validation or verification checkpoint — no smoke test that the endpoint responds before ramping, no pre-flight check beyond tool installation, no guidance on validating suspicious results. Per the rubric's cap for batch operations without validation, workflow clarity cannot exceed 3. Not 2 because usage, prerequisites, and interpretation of output are clearly laid out.

3 / 5

Progressive Disclosure

The body is well-sectioned (Quick Start, Parameters, Output Format, Prerequisites, Algorithm, Notes) and its one bundle path, 'python3 scripts/http_benchmark.py', refers to a real, matching script in the bundle. Not 5 because the file carries the full human-readable and JSON output format examples plus a redundant Features section — some of this belongs in a separate reference — and not 3 because structure is clear and nothing is buried or wrongly inlined to the point of impeding navigation.

4 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete, with an explicit 'Triggered by' clause covering natural user phrasings. Its only weakness is a handful of missing natural synonyms (e.g., 'stress test'), which keeps trigger term quality at 4 rather than 5.

DimensionReasoningScore

Specificity

The description lists four concrete actions with named tools and metrics — 'Run stepped HTTP load tests with ab/wrk', 'collect p50/p90/p99 latency', 'detect performance inflection points', 'recommend optimal concurrency' — comprehensive coverage matching the top anchor. It uses third person throughout with no padding, so it is not the level below where minor gaps in coverage would appear.

5 / 5

Completeness

Both halves are explicit: a concrete 'what' (stepped load tests, percentile collection, inflection detection, concurrency recommendation) and an explicit 'when' clause — 'Triggered by requests like…' with concrete example phrases and keyword mentions. Not 4 because the trigger guidance is fully explicit rather than merely serviceable.

5 / 5

Trigger Term Quality

Strong natural trigger phrases ('load test this URL', 'benchmark my API', 'find the max concurrency', 'p99 latency', 'throughput saturation', 'capacity planning') cover the space well, but common synonyms like 'stress test' or 'how many concurrent users' are absent. Not 5 because keyword coverage has a few natural gaps; not 3 because many phrases a user would actually say are present.

4 / 5

Distinctiveness Conflict Risk

It occupies a clear HTTP load-testing niche with metric-specific triggers (p99 latency, throughput saturation, capacity planning) that competing general benchmarking or profiling skills would not claim. Minimal conflict risk, matching the top anchor; it is not 4 because the triggers are not merely 'mostly distinct' but sharply differentiated.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
zebbern/claude-code-guide
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.