CtrlK
BlogDocsLog inGet started
Tessl Logo

writing-visual-tests

Use when a developer wants to add a visual regression (screenshot) test for an Opik UI page or panel — e.g. "add a visual test for the trace sidebar", "screenshot each tab of the dataset panel", "visual regression test for the new empty state". Covers the page-object pattern, seeding via the test-helper-service, unique screenshot naming, per-test masking, baseline (re)generation, and the local-run-until-stable loop in tests_end_to_end/visual-tests/.

76

Quality

95%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced workflow skill with strong validation and feedback loops for a fragile visual-testing process. Its main weakness is conciseness — the anti-patterns table and narrative asides restate prior content — and the lack of any progressive file split for a skill this large.

Suggestions

Collapse or remove the Anti-patterns table: every row restates a rule already detailed in the body, so it costs tokens for redundancy; keep the body rules and drop the table (or replace it with a one-line pointer to the relevant sections).

Trim the meta-narrative asides (e.g. "This isn't hypothetical — it happened during this skill's own authoring session") which add length without operational value.

Consider extracting the long Cleanup checklist table and the Masks deep-dive into a references/ file linked from the body, so the SKILL.md overview stays leaner while keeping the detail one level deep.

DimensionReasoningScore

Conciseness

Mostly efficient and project-specific, but the Anti-patterns table largely restates the body sections and the authoring-session anecdote ("This isn't hypothetical — it happened during this skill's own authoring session") is padding that could be trimmed.

4 / 5

Actionability

Fully executable guidance throughout: copy-paste page-object and mask code, concrete bash commands (docker ps, lsof -ti:5555, the playwright invocations), and exact file paths covering the common cases.

5 / 5

Workflow Clarity

"The loop" is a clearly sequenced 7-step process with explicit validation checkpoints (stress-test 3×, open the *-diff.png before changing anything), feedback loops (regenerate baseline then re-stress), and a cleanup checklist for the destructive/batch operations.

5 / 5

Progressive Disclosure

Well-organized with clear section headers and a navigable structure, but it is a single monolithic ~350-line file with no bundle files or one-level-deep references; sections like the cleanup checklist and masks deep-dive could be split out.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that names concrete capabilities, supplies natural trigger phrases with synonyms, and carves out a distinct niche. It cleanly answers both what the skill does and when to invoke it.

DimensionReasoningScore

Specificity

Lists multiple concrete capability areas — "page-object pattern, seeding via the test-helper-service, unique screenshot naming, per-test masking, baseline (re)generation, and the local-run-until-stable loop" — giving comprehensive coverage rather than vague abstraction.

5 / 5

Completeness

Explicitly answers both: "Covers the page-object pattern, seeding …" (what) and "Use when a developer wants to add a visual regression (screenshot) test …" (when), with concrete trigger examples.

5 / 5

Trigger Term Quality

Includes natural user phrases with synonyms — "add a visual test", "screenshot each tab", "visual regression test" — covering the ways a developer would actually request this.

5 / 5

Distinctiveness Conflict Risk

Scoped to a clear niche — Opik UI visual regression tests in tests_end_to_end/visual-tests/ — and explicitly distinguishes itself from the functional e2e suite, minimizing wrong-skill triggers.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
comet-ml/opik
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.