CtrlK
BlogDocsLog inGet started
Tessl Logo

explore-feature

Use when a developer wants an e2e test covering a change they just made — e.g. "explore this feature", "add a test for my PR", "cover the feature in PR

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality orchestration skill: an unambiguous gated workflow with explicit validation and stop conditions, entirely copy-paste-ready commands, and dense project-specific knowledge with no filler. The two deductions are minor: some nested case-law prose could be tightened, and the deepest operational recipes (worktree/socat bridging, skip-check edge cases) are inlined in SKILL.md rather than split into a reference file.

Suggestions

Extract the worktree/socat backend-bridging recipe (Phase 2, lines 160-173) into a references/local-run-gate.md file and keep a one-line trigger in SKILL.md ('FE-from-source against a prebuilt backend: see references/local-run-gate.md') — it is deep troubleshooting detail only needed in an uncommon stack configuration.

Restructure the Skip-check paragraph (lines 117-135) as a short decision table (change type → cover/skip action) instead of nested parenthetical case analysis; the current prose requires multiple re-reads to extract the decision rule.

Similarly consider moving the three 'false or unbuildable test' piloting lessons (lines 88-108) to a reference file, keeping the one-line rule of each inline — SKILL.md would then read as a lean overview with detail one level deep.

DimensionReasoningScore

Conciseness

Every section carries project-specific earned knowledge (port proxies, worktree port-hashing, tag_lint as a CI gate) rather than concepts Claude already knows — git/docker/Playwright mechanics are used, never taught. It sits at the 4 anchor ("efficient; minor instances of over-explanation that could be trimmed") rather than 5 because the nested case-law prose in the Skip-check paragraph (lines 117-135) and the socat worktree gotcha are denser than needed and could be tightened.

4 / 5

Actionability

Fully executable, copy-paste-ready commands throughout: the four git diff variants, `curl <baseUrl>/api/is-alive/ver`, the complete `docker run -d ... alpine/socat ...` bridge with the `VITE_DEV_PORT=5174` start line, `tag_lint.py` invocation, `npx playwright test tests/<area>/ --reporter=list` and `npx tsc --noEmit`. Matches the 5 anchor ("copy-paste ready code or commands"); 4 would require gaps in coverage of the common cases, and there are none.

5 / 5

Workflow Clarity

Four explicitly gated phases with a digraph of the loop, hard validation checkpoints (tag_lint `0 problem(s)`, green run, tsc clean, taxonomy updated), explicit stop conditions ("If all four are empty... say so and stop"; skip-with-a-note rules), and feedback routing ("QA owns this skill... the fix lands in this skill's files"). This is the 5 anchor — clear sequence, explicit validation, error-recovery guidance — not 4, which allows missing checkpoints.

5 / 5

Progressive Disclosure

The body is well-sectioned (phases, headers, tables) and demonstrates a genuine one-level-deep deferral — "Don't restate them here; read `.agents/skills/writing-e2e-tests/conventions.md` if you need them" — instead of duplicating another skill's conventions. No bundle files exist (references/, scripts/, assets/ are absent), so everything lives in SKILL.md; the ~15-line worktree/socat bridging recipe and the skip-check case analysis are inlined detail that would fit a reference file. That places it at the 4 anchor ("good structure; most content appropriately placed; minor organization gaps") rather than 5.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit 'Use when' with several realistic trigger phrasings, a concrete multi-step capability statement in third person, and a clear niche. The only weakness is mild ambiguity against its sibling authoring skill, since a request phrased as 'add a test for my PR' could plausibly route to either.

Suggestions

Sharpen the boundary with the sibling skill in the description, e.g. '...when the dev wants a NEW test scoped to their change; for edits to existing e2e specs, use writing-e2e-tests directly' — this would remove the residual trigger ambiguity and lift distinctiveness to 5.

DimensionReasoningScore

Specificity

Concrete actions are enumerated end-to-end: "Reads a ticket + changed code to work out the one flow worth covering, then delegates authoring to writing-e2e-tests, which writes a normal tiered spec under tests/<area>/" — that is comprehensive coverage of this skill's scope, not just 1-2 actions, so it matches the 5 anchor rather than the 4.

5 / 5

Completeness

Both halves are explicit: what (reads ticket + changed code, resolves the one flow, delegates to writing-e2e-tests, produces a tiered spec under tests/<area>/) and when ("Use when a developer wants an e2e test covering a change they just made") with concrete trigger phrases. This is the 5 anchor verbatim in structure; 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Multiple natural user phrasings with synonym coverage: "explore this feature", "add a test for my PR", "cover the feature in PR #7303", "test my branch", plus "e2e test". This matches the 5 anchor (comprehensive natural terms incl. variations), not 4 ("a few natural terms missing") — the common ways a dev would ask for this are all present.

5 / 5

Distinctiveness Conflict Risk

The description carves a clear niche (thin orchestrator that resolves scope then delegates) and names the sibling skill it delegates to, but triggers like "add a test for my PR" carry minor overlap risk with the closely related writing-e2e-tests authoring skill itself. This fits the 4 anchor ("mostly distinct; minor overlap risk with closely related skills"); 5 would require no such sibling ambiguity.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
comet-ml/opik
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.