CtrlK
BlogDocsLog inGet started
Tessl Logo

author-e2e-tests

Use when writing or maintaining Playwright e2e tests for Positron -- new test files, test cases, test infrastructure, or performance/metric tests. For a named test that is already failing or flaking, in CI or on your machine, use debug-e2e-test instead.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, dense, project-specific skill: an explicit starting workflow, mandatory requirements with rationale, executable templates and commands, and a well-organized one-level-deep reference bundle. The only trims worth making are the generic test-philosophy section and standard Playwright CLI examples, and some inline quick-reference content that duplicates the bundled references.

DimensionReasoningScore

Conciseness

Nearly all content is Positron-specific knowledge Claude cannot know (the `_test.setup` import rule, `suiteId`, fixture table, tag enums, the `sendEnterKey()` ~600ms timing trap), so tokens largely earn their place. Minor over-explanation remains: the 'Philosophy' section reiterates generic good-test principles Claude already knows, and the 'Running Tests' block includes standard Playwright CLI knowledge (`--headed`, `--debug`, `show-report`), fitting 'efficient; minor instances of over-explanation that could be trimmed'.

4 / 5

Actionability

The skill is copy-paste ready throughout: a complete executable file template, a fixture table with real signatures (`await sessions.start('python')`), exact import and tag rules with source file paths (`test/e2e/infra/test-runner/test-tags.ts`), and concrete run commands. The performance section even quantifies the pitfall ("inflates every measurement by ~600ms") with a precise workaround.

5 / 5

Workflow Clarity

The workflow is explicitly sequenced: 'Start Here: Read a Neighbor Test' gives the entry action, then the mandatory structure, fixtures, page objects, and assertions, with verification steps baked in ("Never guess or paraphrase a method name -- copy it from the source file", "check its source for waitForTimeout") and an error-recovery ladder in 'Getting Help' ending in a --debug loop and hand-off to debug-e2e-test.

5 / 5

Progressive Disclosure

Six real, one-level-deep reference files exist and each is clearly signaled with its scope (e.g. "references/assertions.md - Retry-mechanism choice and selector priority"), matching 'good structure; references mostly clear'. It falls short of the 5 anchor because the body inlines condensed duplicates of reference content (the fixture table mirrors references/fixtures.md and the 'Common Mistakes' list mirrors references/common-mistakes.md) rather than pointing to them the way the form-filling example does.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-crafted description: explicit 'Use when' triggers, concrete enumerated scope, tight product/tool scoping, and an explicit hand-off to a sibling skill for the debugging case. Minor room to add common synonyms like 'end-to-end' or spec-file extensions.

DimensionReasoningScore

Specificity

The description names the concrete activities "writing or maintaining Playwright e2e tests" and enumerates specific artifacts ("new test files, test cases, test infrastructure, or performance/metric tests"), matching the anchor for several specific actions with minor gaps. It does not reach 5 because the actions stay at the task level rather than enumerating the concrete operations the skill covers (e.g. fixtures, page objects, assertions).

4 / 5

Completeness

It opens with an explicit trigger clause ("Use when writing or maintaining Playwright e2e tests for Positron") and scopes what is covered ("new test files, test cases, test infrastructure, or performance/metric tests"), clearly answering both what and when. It additionally disambiguates the negative case ("For a named test that is already failing or flaking... use debug-e2e-test instead"), making the when fully explicit.

5 / 5

Trigger Term Quality

Strong natural terms a user would say: "Playwright e2e tests", "Positron", "test files", "test cases", "flaking", "in CI", "performance/metric tests". It misses a few common variations such as "end-to-end", "spec", or file extensions like ".spec.ts", keeping it at 'good keyword coverage' rather than comprehensive.

4 / 5

Distinctiveness Conflict Risk

The niche is tightly defined (Positron + Playwright e2e authoring) and it explicitly routes the adjacent failing/flaky-test case to a different skill ("use debug-e2e-test instead"), minimizing overlap risk. It is at least as distinct as the 5-anchor example, which has no such explicit de-confliction.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.