CtrlK
BlogDocsLog inGet started
Tessl Logo

writing-e2e-tests

Use when a developer wants to add, write, or create an end-to-end test for an Opik feature, page, or branch — e.g. "add an e2e test for the experiments comparison page", "write a test for the feature I just built", "e2e test for this branch", "cover the dataset items flow with a test". Runs the full loop in tests_end_to_end/e2e/ — analyze the feature and frontend code, explore the live UI with the Playwright MCP, write the Page Object Model + spec, and run it locally until green.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable skill body: concrete commands everywhere, a gated loop with real validation checkpoints, and disciplined references to conventions.md. The only costs are a redundant graphviz loop diagram and inlined docker-rebuild plumbing that could live in the referenced materials.

Suggestions

Drop the graphviz digraph (or shrink it to a one-line flow list) — it restates the five step headers that follow it verbatim.

Move the FE rebuild / container-network troubleshooting block into conventions.md (or a dedicated reference) and keep a one-line pointer plus the non-negotiable 'rebuild before running' reminder in the step.

Ensure conventions.md actually ships alongside SKILL.md in the bundle, since the body depends on it for selector rules, fixture shapes, and the tag taxonomy.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence — no library tutorials, no filler — but the 18-line graphviz 'digraph' duplicates the numbered step headers that follow it, and the docker-rebuild/network-troubleshooting detail (~25 lines) could be trimmed or moved. Not 5: those redundant and peripheral blocks are tokens that don't earn their place; not 3: the rest is uniformly lean with zero concept re-teaching.

4 / 5

Actionability

Fully executable throughout: `npx playwright test tests/<feature>/<name>.spec.ts --reporter=list`, the complete docker compose rebuild sequence, `python3 tests_end_to_end/coverage/tag_lint.py ...`, `npx tsc --noEmit`, and the exact `~/.opik.config` heredoc including the backup step. Not 4: commands are copy-paste ready and the common cases (build image, recreate container, verify testid, restore config) are each covered end-to-end.

5 / 5

Workflow Clarity

A five-step sequence with two explicit gates, an explicit feedback loop ('Run until green' → read the failure trace → fix → re-run), a safety check before any seeding, and a three-item completion checklist with the rationale for each. Not 4: validation checkpoints are present at every risky point, including the destructive-capable config rewrite and the taxonomy/tsc checks that 'never surface in a Playwright run'.

5 / 5

Progressive Disclosure

Good structure: the body is a clear overview that repeatedly and purposefully signals one external file, conventions.md, at the exact points it matters ('Read conventions.md before writing any POM or spec', the Tags section), plus pointers to taxonomy.yaml and the playwright-pom-discovery skill. Not 5: the referenced conventions.md is not present in this bundle (no references/ directory) so the split cannot be verified, and operational detail like the FE rebuild/network-fix section is inlined in the main file rather than separated; not 3: references are clearly signaled at point-of-need rather than buried, and the main file stays an overview.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person, explicit 'Use when' triggers with verbatim user phrasings, a complete concrete action list, and a well-delineated niche. Nothing material is missing or padded.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'analyze the feature and frontend code, explore the live UI with the Playwright MCP, write the Page Object Model + spec, and run it locally until green' — which comprehensively covers the skill's loop. Not 4: there are no coverage gaps; every phase of the workflow is named.

5 / 5

Completeness

It explicitly answers both questions: 'Runs the full loop in tests_end_to_end/e2e/ — analyze... write... run it locally until green' (what) and 'Use when a developer wants to add, write, or create an end-to-end test for an Opik feature, page, or branch' (when) with concrete trigger phrases. Not 4: the 'when' is fully explicit with example utterances, not just implied.

5 / 5

Trigger Term Quality

It covers natural variations users would say: 'add, write, or create an end-to-end test', plus four verbatim example phrasings ('add an e2e test for the experiments comparison page', 'write a test for the feature I just built', 'e2e test for this branch'). Not 4: it includes both the full term and the 'e2e' shorthand with realistic sentence-level triggers, leaving essentially no common variation missing.

5 / 5

Distinctiveness Conflict Risk

It occupies a clear niche — Opik's E2E suite, `tests_end_to_end/e2e/`, Playwright Page Object Model — with triggers anchored to that domain. Not 4: the product name and suite path make overlap with generic test-writing skills negligible.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 missing

Warning

Total

15

/

16

Passed

Repository
comet-ml/opik
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.