CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-agents

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Designed to stress-test skill injection for complex, multi-system builds.

55

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/benchmark-agents/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a highly actionable, well-sequenced operational skill with exact commands, timing guidance, explicit verification checkpoints, and a genuine feedback loop. Its weaknesses are repetition and embedded version-history that inflate token cost, and a monolithic single-file structure that inlines scenario and issue-history tables which belong in separate reference files.

Suggestions

Split the Scenario Table and the Common Issues version-history table into reference files (e.g. references/scenarios.md, references/issue-history.md) and keep a brief summary in SKILL.md, which would also remove the time-sensitive version numbers from the always-loaded context.

Deduplicate the DO-NOT section against the "How Evals Work" rationale (each rule states its reason once) and replace the three repeated wezterm spawn commands in the parallel section with a single parameterized loop.

Make the abbreviated blocks fully executable: expand `--cwd .../tarot-deck-$TS`, show a complete example prompt in place of `'<PROMPT>'` / `'...'`, and give a command to resolve the session-id used in `~/.claude/debug/<session-id>.txt`.

DimensionReasoningScore

Conciseness

Mostly efficient commands, grep patterns, and tables, but there is real padding: the DO-NOT section restates the hook-firing rationale already given in "How Evals Work", the parallel-launch section repeats the full wezterm spawn command three times, and the 14-row issue table embeds time-sensitive version numbers (v0.8.0–v0.9.9) outside any "old patterns" section. Not 4 because these repetitions and the historical version log are unnecessary tokens; not 2 because there is no explaining of concepts Claude already knows and the bulk is actionable command content.

3 / 5

Actionability

Setup, launch, monitoring, and verification all give copy-paste-ready Bash with exact flags, timings ("wait ~25s"), and concrete checks like `ls "$CLAIMDIR/workflow" && echo "YES" || echo "NO"`. Not 5 because several blocks are abbreviated rather than executable: `wezterm cli spawn --cwd .../tarot-deck-$TS`, prompts shown as `'<PROMPT>'` and `'...'`, and `~/.claude/debug/<session-id>.txt` placeholders the user must reconstruct. Not 3 because the gaps are minor against an overwhelmingly concrete body.

4 / 5

Workflow Clarity

The full eval loop (setup → launch → monitor → verify → fix → release → repeat) is explicitly sequenced with numbered steps, an explicit timing checkpoint ("wait ~25s for SessionStart hooks"), a verification section with concrete checks before agent-browser verification, a coverage-report checklist, and a fix→re-run feedback loop ("Release → Eval Loop" steps 1–8). The destructive `rm -rf` cleanup occurs only after verification and reporting. Not 4 because both checkpoints and error-recovery loops are explicit throughout.

5 / 5

Progressive Disclosure

Section headers are clear and well-ordered, but no bundle files exist (no references/, scripts/, assets/) and the ~310-line body inlines content that clearly belongs in separate files — the 12-scenario table, the 14-row issue-history table, and the verification grep sets would all work better as one-level-deep reference files or scripts. Not 4 because significant content that should be separate is inline with no references at all; not 2 because the structure and navigation within the file are good, not minimal.

3 / 5

Total

15

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly identifies a distinctive niche (Vercel plugin benchmark/eval scenarios) with a concrete feature list, but it describes the product being tested rather than what the skill does, and it contains no "Use when..." trigger guidance. It reads as marketing copy with jargon-heavy phrases rather than a user-voice trigger description.

Suggestions

Add an explicit trigger clause, e.g. "Use when running evals or benchmark sessions for the Vercel plugin, verifying skill injection, or producing coverage reports."

State the skill's concrete actions in third person — launch Claude Code sessions in WezTerm, monitor skill-injection claims and hook debug logs, verify generated code patterns, write a coverage report — instead of "push" and "stress-test" abstractions.

Replace promotional language ("cutting-edge", "stress-test", "complex, multi-system builds") with natural user-voice keywords like "run evals", "test plugin injection", "check hook coverage".

DimensionReasoningScore

Specificity

Names the domain and enumerates concrete platform features ("Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox"), but the stated actions are limited to vague verbs like "push" and "stress-test skill injection" — it never states what the skill actually does (launch sessions, monitor hooks, produce coverage reports). Not 4 because several specific actions are missing; not 2 because the domain and feature list are concrete, exceeding minimal/generic.

3 / 5

Completeness

The "what" is clear (benchmark scenarios stress-testing Vercel platform features and skill injection), but there is no "Use when..." clause or any equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Not 4 because the "when" is entirely absent rather than weakly implied; not 2 because the "what" is explicit and specific.

3 / 5

Trigger Term Quality

"benchmark scenarios", "AI agent", and "multi-agent" are terms a user might say, but the rest is product jargon ("skill injection", "multi-system builds", "stress-test") and there are no synonyms or phrasings a user would naturally type when they need this. Not 4 because common natural variations (e.g. "run evals", "test the plugin") are absent; not 2 because more than one relevant keyword is present.

3 / 5

Distinctiveness Conflict Risk

The Vercel-plugin benchmarking niche with named subsystems (Workflow SDK, AI Gateway, Chat SDK) is fairly distinct from generic Vercel or general agent skills. Not 5 because phrases like "multi-agent orchestration" and "AI agent" could overlap with unrelated agent-building skills; not 3 because the specific platform-feature list keeps it mostly distinguishable.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
vercel/vercel-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.