CtrlK
BlogDocsLog inGet started
Tessl Logo

storybook-startup-benchmark

Measure Storybook startup time from spawning `storybook dev` until the first story renders in the browser. Use when the user asks about Storybook boot time, server-ready timing, first story render timing, startup regressions, benchmarking with repeat runs, or comparing Storybook versions or feature flags.

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-crafted, dense skill body: precise measurement boundaries, concrete implementation details, explicit validation and cleanup steps, and genuinely useful domain knowledge (warm vs cold runs, iframe.html pitfall, server/browser/total attribution). Its weaknesses are moderate redundancy across three overlapping step restatements and the absence of a complete executable script plus hang/timeout error handling.

Suggestions

Consolidate the overlapping step sequences in 'Quick Start', 'Measurement Rules', and 'Harness behavior' into one canonical numbered sequence, keeping the other sections for defaults and rationale only — this removes ~20 lines of repetition.

Add explicit error-recovery guidance for the harness hanging (e.g., timeout on the preview-side signal with a diagnostic message) to complete the feedback loop alongside the existing stale-server check.

Consider moving the full summary-stats JSON block and detailed harness spec into a one-level-deep reference file (e.g., references/summary-format.md) to keep SKILL.md as a tighter overview.

DimensionReasoningScore

Conciseness

Sections are lean, imperative, and free of concept-explanation padding, but the Quick Start steps, 'Measurement Rules', and 'Harness behavior' restate the same spawn→wait→launch→signal→print sequence three times, which could be tightened. Fits 4 (efficient with minor trimmable overlap) rather than 5 (every token earns its place).

4 / 5

Actionability

Concrete flags, APIs, and schemas are given (`storybook dev --no-open`, `requestAnimationFrame()`, `window.__sbStartupBenchmark`, `performance.mark('sb:first-story-rendered')`, explicit JSON payload and summary formats), but no complete executable harness script is provided — a gap partly justified by the 'reuse an existing benchmark script if possible' instruction, which places it at 4 rather than 5 or 3.

4 / 5

Workflow Clarity

The 8-step harness sequence includes explicit validation (fail fast if the port is in use), cleanup (kill the full process group), and a feedback loop (unrealistically low `server.average` → check for a stale server), plus a pitfalls checklist. It falls short of 5 only because there is no timeout or no-signal-received error handling for a harness that could hang waiting on the preview-side signal.

4 / 5

Progressive Disclosure

The single body file is well-organized with clear headers and appropriately self-contained guidance, but at ~157 lines the detailed harness spec and full summary JSON block could be offloaded to a one-level-deep reference file, which matches the 4 anchor (good structure, minor organization gaps) rather than 5.

4 / 5

Total

16

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: a single precisely-bounded capability stated in imperative voice, paired with an explicit and thorough 'Use when' trigger clause rich in natural synonyms. The only minor limitation is that the what-clause carries one core action with facets like repeat-run stats left to the trigger clause.

DimensionReasoningScore

Specificity

"Measure Storybook startup time from spawning `storybook dev` until the first story renders in the browser" defines one precise, boundary-explicit action, and supporting facets (repeat runs, version/flag comparison) appear only in the trigger clause, leaving minor coverage gaps that fit the 4 anchor rather than the comprehensive 5.

4 / 5

Completeness

It explicitly answers both what ("Measure Storybook startup time from spawning `storybook dev` until the first story renders") and when ("Use when the user asks about Storybook boot time, server-ready timing...") with concrete trigger phrases, directly matching the 5 anchor exemplar.

5 / 5

Trigger Term Quality

The trigger clause covers a comprehensive set of natural phrasings users would say — "Storybook boot time", "server-ready timing", "first story render timing", "startup regressions", "benchmarking with repeat runs", "comparing Storybook versions or feature flags" — with effective synonym coverage and no obvious missing variation for this domain.

5 / 5

Distinctiveness Conflict Risk

"Storybook startup benchmark" is a clearly delimited niche with named tooling and distinct trigger terms; minimal risk of firing for any unrelated skill.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

13

/

16

Passed

Repository
storybookjs/storybook
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.