CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark

Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/benchmark/SKILL.md

The canonical home for this skill is tdg-personal/benchmark

SKILL.md
Quality
Evals
Security

Quality

Content

72%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is concise and well-structured with concrete commands and targets, but Modes 1–3 lack executable measurement code and no validation checkpoints exist for the batch operations.

Suggestions

Replace the descriptive Mode 1–3 step lists with executable guidance — concrete browser-MCP calls for Web Vitals, a runnable script or curl/hey command for API p50/p95/p99, and actual build-timing commands — so the skill is copy-paste ready.

Add validation/verification checkpoints for the batch operations, e.g., flag anomalous p99 spikes, confirm baseline stability across repeated runs, and re-measure on outlier detection before saving a baseline.

DimensionReasoningScore

Conciseness

The body is lean: it lists Core Web Vitals with numeric targets and resource categories without explaining what they are, and uses compact numbered code blocks per mode — every token earns its place and it assumes Claude's competence.

3 / 3

Actionability

Mode 4 gives concrete commands ('/benchmark baseline', '/benchmark compare') and an output table, but Modes 1–3 are descriptive step lists ('Measure Core Web Vitals', 'Hit each endpoint 100 times') with no executable measurement code or specific browser-MCP calls, so guidance is incomplete.

2 / 3

Workflow Clarity

Steps are clearly sequenced and Mode 4 has a clean before/after flow, but the batch operations ('Hit each endpoint 100 times', '10 concurrent requests') have no validation or verification checkpoints, capping workflow clarity at 2 per the rubric.

2 / 3

Progressive Disclosure

No bundle files exist or are needed; the single SKILL.md is well-organized into clear sections (When to Use, How It Works with four modes, Output, Integration), satisfying the simple-skill allowance for well-organized content.

3 / 3

Total

10

/

12

Passed

Description

57%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinct but loses points for second-person imperative voice, a missing explicit 'Use when...' trigger clause, and thin trigger-term coverage.

Suggestions

Reword in third person (e.g., 'Measures performance baselines, detects regressions before/after PRs, and compares stack alternatives. Use when the user mentions benchmarks, perf, latency, or that something feels slow.') to remove the second-person specificity penalty.

Add an explicit 'Use when...' clause listing natural user phrases ('benchmark performance', 'it feels slow', 'measure regressions', 'compare stacks') to lift completeness to 3.

Broaden trigger terms to include common variations a user would actually say — 'benchmark', 'perf', 'speed', 'latency' — alongside the current vocabulary.

DimensionReasoningScore

Specificity

The description names three concrete actions ('measure performance baselines, detect regressions before/after PRs, and compare stack alternatives'), which would merit a 3, but the imperative 'Use this skill to' is second-person voice and triggers the −1 specificity penalty.

2 / 3

Completeness

The 'what' is clearly stated, but the 'when' is only implied via 'before/after PRs'; there is no explicit 'Use when...' clause, so completeness is capped at 2 per the rubric guidelines.

2 / 3

Trigger Term Quality

Terms like 'performance baselines', 'regressions', and 'before/after PRs' are reasonably natural, but coverage is thin — common variations a user would say ('benchmark', 'perf', 'speed', 'latency', 'it feels slow') are absent.

2 / 3

Distinctiveness Conflict Risk

Performance baseline and regression detection is a clear niche with distinct triggers ('regressions', 'performance baselines', 'compare stack alternatives') unlikely to fire for unrelated skills.

3 / 3

Total

9

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.