CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-sandbox

Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels. Provisions ephemeral microVMs with Claude Code + plugin pre-installed, runs benchmark prompts, extracts hook artifacts, and produces coverage reports.

79

2.09x
Quality

70%

Does it follow best practices?

Impact

92%

2.09x

Average score across 3 eval scenarios

SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/benchmark-sandbox/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable with a clearly sequenced, well-validated 3-phase workflow and hard-won environment specifics that Claude could not know. Its weaknesses are redundancy across overlapping fact sections and a monolithic single-file layout whose only file reference (run-eval.ts) is absent from the bundle.

Suggestions

Consolidate the overlapping "Critical Sandbox Environment Facts", "Key Discoveries", "DO NOT", and "Known Limitations" sections into one canonical facts section, and remove the "Commands" section that duplicates the "Proven Working Script" examples.

Move the score schemas, artifact export layout, and dated "Proven Results" table into a separate reference file (e.g., references/scoring.md or references/results.md) to slim SKILL.md into an overview.

Ship run-eval.ts in the skill's scripts/ directory (or reference its actual location) so the documented invocation path resolves to a real file.

DimensionReasoningScore

Conciseness

The bulk is high-value, non-obvious operational knowledge, but facts repeat across sections (home dir, snapshot behavior, and v2-beta 404 each appear three times) and the "Commands" section duplicates the "Proven Working Script" invocations. Time-stamped "Proven Results (2026-03-10)" and version pins add staleness outside any deprecated section. Not 2 because it never explains concepts Claude already knows and most content earns its place.

3 / 5

Actionability

Fully copy-paste ready throughout: exact bun run invocations, a complete CLI flags table, the scenario JSON format, auth setup commands, executable monitoring snippets, and concrete structured scoring schemas cover the common cases.

5 / 5

Workflow Clarity

The 3-phase pipeline is explicitly sequenced (17-step walkthrough plus per-scenario session flow with per-phase timeouts) with validation checkpoints and feedback loops: haiku scoring after each phase, deploy retries up to 3 with fixes, and verify's fix-and-re-verify loop. The batch-operation cap does not apply since validation is present.

5 / 5

Progressive Disclosure

Section structure with headers and tables is good, but everything is inlined in one ~385-line file — score schemas, artifact layouts, and dated results data that belong in separate reference files — and the sole referenced script (run-eval.ts) does not exist in the bundle. Not 4 because the broken script reference and monolithic inline content exceed minor organization gaps.

3 / 5

Total

16

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinctive with four concrete capabilities, but it omits any trigger guidance for when Claude should use it, which limits completeness and natural keyword coverage. Adding a "Use when..." clause with user-facing trigger phrases would raise it substantially.

Suggestions

Append a trigger clause such as "Use when running vercel-plugin benchmarks, evaluating plugin skill coverage, or when the user asks to run eval scenarios in sandboxes rather than locally."

Include natural synonym phrasings users would say (e.g., "benchmark run", "plugin eval", "skill coverage report") to broaden trigger-term coverage.

Replace the generic "runs benchmark prompts" with a more concrete action like "runs 3-phase build/verify/deploy benchmark scenarios".

DimensionReasoningScore

Specificity

Lists several concrete actions — "Provisions ephemeral microVMs", "extracts hook artifacts", "produces coverage reports" — matching the several-specific-actions anchor with minor coverage gaps ("runs benchmark prompts" is generic and the 3-phase pipeline is compressed).

4 / 5

Completeness

The "what" is clear and multi-part, but there is no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Not 2 because the "what" is specific rather than vague.

3 / 5

Trigger Term Quality

Relevant domain keywords exist ("eval scenarios", "benchmark prompts", "coverage reports") but common natural variations and synonyms a user might say are missing. Not 4 because keyword coverage is thin beyond the exact domain terms.

3 / 5

Distinctiveness Conflict Risk

"vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels" carves out a clear niche with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
vercel/vercel-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.