CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-e2e

End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill injection, launches dev servers, verifies everything works, analyzes conversation logs, and produces an improvement report for overnight self-improvement loops.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/benchmark-e2e/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a strong, executable skill document: concrete commands, precise contracts, clear sequencing with abort-on-failure semantics, and a well-documented feedback loop. The only notable gap is the absence of any reference files despite referencing scripts/benchmark-e2e.ts, leaving all detail inlined in the single file.

DimensionReasoningScore

Conciseness

The body is efficient and information-dense (tables, typed contracts, commands), with only minor trimmable material — the intro paragraph restates the description and lines like "Wake up to reports showing exactly what improved and what still needs work" are motivational padding.

4 / 5

Actionability

Commands are copy-paste ready ("bun run scripts/benchmark-e2e.ts --quick", the "while true" automation loop, the "rm -rf" cleanup), the flags table and typed interfaces (BenchmarkRunManifest, ReportJson) are concrete, and the self-improvement cycle gives exact steps including using "suggestedPatterns entries (copy-pasteable YAML)".

5 / 5

Workflow Clarity

The four pipeline stages are clearly sequenced with an explicit "aborting on failure" checkpoint and a documented run → read gaps → apply fixes → re-run feedback loop, plus error/abort event examples; guidance for recovering from a failed stage mid-run (beyond aborting) is absent, keeping it just below the top anchor.

4 / 5

Progressive Disclosure

Sections are well-organized and navigation is easy, and validation is naturally built into the verify stage so the batch-operation cap does not apply. However, the ~140-line body is fully inline with no bundle files in the bundle (references/, scripts/, assets/ are absent), and inlined content like the 9-row Prompt Table could live in a one-level-deep reference file.

4 / 5

Total

17

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description does a good job of enumerating concrete pipeline capabilities and is anchored to a specific, distinguishable niche. Its main weakness is the complete absence of "when to use" trigger guidance and natural user-facing trigger terms.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when you want to test whether the vercel-plugin injects the right skills, or when running the overnight benchmark/self-improvement loop.'

Include natural trigger phrases users would actually say, such as 'run the e2e benchmarks', 'check skill injection coverage', or 'generate a plugin improvement report', to improve trigger term quality.

Replace the vague clause 'verifies everything works' with a concrete action like 'verifies dev servers return 200 with non-empty HTML' to push specificity toward comprehensive coverage.

DimensionReasoningScore

Specificity

Quotes like "Runs realistic projects through skill injection", "launches dev servers", "analyzes conversation logs", and "produces an improvement report" list several concrete actions, but "verifies everything works" is generic, leaving minor gaps that keep it below the comprehensive anchor.

4 / 5

Completeness

The description clearly answers "what" with multiple concrete pipeline actions, but contains no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Domain keywords such as "benchmark suite", "dev servers", and "conversation logs" are relevant but jargon-leaning, and the natural phrases a user would say to invoke this skill (e.g., "run the benchmarks", "check the e2e suite") are missing.

3 / 5

Distinctiveness Conflict Risk

"End-to-end benchmark suite for vercel-plugin" pins a clear niche tied to a specific plugin with minimal conflict risk, though it lacks distinct trigger phrases and could marginally overlap with generic testing or benchmark skills.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

15

/

16

Passed

Repository
vercel/vercel-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.