CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/visual-baseline-conventions

Reference catalog for visual regression coverage decisions - which Storybook stories or pages get baselines, how to choose breakpoints, when to mask vs adjust threshold, when to add or remove a baseline, and a decision matrix for picking among Percy / Chromatic / Playwright / Storybook test-runner. Use when designing visual coverage for a new project or auditing an existing baseline set.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, well-structured reference catalog with concrete values, decision matrices, and explicit anti-patterns. Its main weakness is mild redundancy between the inline 'Common anti-patterns' table and earlier per-topic callouts.

Suggestions

Consolidate the 'Common anti-patterns' table with the inline 'Anti-pattern' callouts in each section to remove the restated duplication and tighten conciseness.

Consider extracting the detailed engine-selection and breakpoint/masking tables into a references/ file so SKILL.md stays a concise overview with one-level-deep pointers.

Add a short explicit validation checkpoint for baseline updates (e.g. confirm the snapshot diff accompanies the code diff in the same PR review) to raise workflow clarity.

DimensionReasoningScore

Conciseness

Mostly lean, table-driven, and free of generic-concept padding, but the 'Common anti-patterns' table restates points already made in earlier sections and a few rationale sentences could be trimmed.

4 / 5

Actionability

Highly actionable for a catalog skill: exact breakpoint widths (375/768/1280/1920 px), concrete threshold ranges (maxDiffPixels 50-200, threshold 0.2->0.3), specific tool flags, and worked naming examples — absence of code is appropriate here.

5 / 5

Workflow Clarity

Decision flows are clearly ordered (e.g. 'wait -> mask -> threshold' preference, baseline add/remove/update lifecycle, warn->block promotion after ~2 weeks) with explicit gating guidance, but there are no validate->fix->retry checkpoints because the skill drives decisions rather than performing operations.

4 / 5

Progressive Disclosure

Well-organized with clear section headers and one-level signaled references to sibling skills (e.g. playwright-snapshots references/responsive-breakpoints.md); no bundle files exist so everything is inline, and the duplicated anti-patterns table is a minor organization gap.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it states concrete capabilities, names natural trigger terms and tool synonyms, answers both what and when, and explicitly distinguishes itself from sibling engine skills. No improvements needed.

DimensionReasoningScore

Specificity

Lists multiple concrete decisions — 'which Storybook stories or pages get baselines, how to choose breakpoints, when to mask vs adjust threshold, when to add or remove a baseline, and a decision matrix' — comprehensively covering the catalog's scope.

5 / 5

Completeness

Explicitly answers both what (the catalog of coverage decisions) and when ('Use when designing visual coverage for a new project or auditing an existing baseline set') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural domain terms and synonyms a user would actually say — 'visual regression coverage', 'Storybook stories', 'breakpoints', 'baselines', plus the full set of tool names 'Percy / Chromatic / Playwright / Storybook test-runner'.

5 / 5

Distinctiveness Conflict Risk

Carves a clear niche — engine-agnostic 'which baselines and where' decisions — distinct from the named engine-specific skills, minimizing wrong-skill triggering.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Reviewed

Table of Contents