CtrlK
BlogDocsLog inGet started
Tessl Logo

browser-automation

Build reliable browser checks using observed UI state, semantic locators, bounded waits, isolated test data and explicit outcome verification.

53

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/browser-automation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, example-driven skill body with exemplary progressive disclosure and mostly executable guidance. Its weaknesses are inline changelog metadata and some over-dense prose that waste tokens, and the absence of an explicitly numbered, checkpointed workflow despite the skill's own warnings about side-effect duplication.

Suggestions

Remove the maintainer-changelog sentence and the modification date from the body (or move it to a deprecated/changelog note); it is time-sensitive information that penalizes token efficiency.

Add a short numbered workflow with explicit validation checkpoints (e.g. 1. inspect page state → 2. choose semantic locators → 3. bounded wait → 4. act → 5. verify outcome/state before any retry), especially given the side-effect-duplication risk the skill itself flags.

Rewrite the over-compressed sentences in the limitations and usability notes into plain, direct statements so they read as instructions rather than riddles.

DimensionReasoningScore

Conciseness

Mostly efficient — the code blocks are lean and comments are minimal — but the body carries time-sensitive changelog prose ("Modified by AAS maintainers on 2026-09-05: removed unverified comparisons and bypass defaults...") that earns no tokens for execution, plus some cryptic, over-compressed sentences ("A locked desktop leaves interactive verification pending; unit tests and headless probes are separate evidence"). This sits between the 3 and 4 anchors: real unnecessary explanation exists but is limited.

3 / 5

Actionability

The body provides concrete, executable Playwright snippets — getByRole/getByLabel/getByTestId examples with good and bad contrasts, and a complete isolated-context test — plus specific directives like "register the download event before clicking". It misses 5 because the end-to-end procedure lives in the reference file and some guidance (e.g. the worked example) is described rather than given as runnable steps.

4 / 5

Workflow Clarity

The worked example implies a sequence (observe the control → register the download event → click → inspect the JSON → confirm invalidation) and the limitations mention verification, but there is no numbered workflow with explicit validation checkpoints in the body. For a skill that explicitly warns about duplicating side effects (checkout, deletion, messaging) that is a real checkpoint gap, matching the 3 anchor; it is above 2 because the sequence and verification requirements are at least stated.

3 / 5

Progressive Disclosure

The body is a genuine overview — activation guidance, contrasting examples, when-to-use, worked example, limitations — with a single clearly-signaled, one-level-deep reference: "Read [the detailed guide](references/detailed-guide.md) before executing this skill", including guidance on partial vs. full reads. The bundle confirms the reference is real and contains no nested references, so navigation is easy and appropriately split.

5 / 5

Total

15

/

20

Passed

Description

55%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description gives a concrete, jargon-accurate picture of what the skill does, but it is written for an expert reader rather than as a trigger surface: it entirely lacks a 'when to use' clause and the natural keywords (tool names, e2e/UI-test phrasings) that would let a user or model select it reliably. Adding an explicit 'Use when...' sentence with those terms would raise both completeness and trigger quality.

Suggestions

Append a trigger clause, e.g. "Use when verifying browser workflows, debugging UI timing or flaky end-to-end tests, or automating pages with Playwright, Puppeteer, or Selenium."

Include the natural tool names users actually say (Playwright, Puppeteer, Selenium, e2e tests, UI tests) so trigger_term_quality reaches comprehensive coverage.

State the 'when' boundary explicitly (fresh automation tests vs. interacting with an already-authenticated browser) to reduce overlap with adjacent automation skills.

DimensionReasoningScore

Specificity

The description lists several concrete technical elements — "semantic locators, bounded waits, isolated test data and explicit outcome verification" — grounded in one clear action ("Build reliable browser checks"). It falls short of the 5 anchor because coverage is technique-listing rather than a comprehensive set of distinct actions, and ahead of 3 because more than 1-2 concrete specifics are given.

4 / 5

Completeness

The 'what' is clear (build reliable browser checks using specific techniques), but there is no 'Use when...' clause or equivalent trigger guidance anywhere in the description, which caps completeness at 3 per the judging guidelines. It is above 2 because the 'what' is concrete, not vague.

3 / 5

Trigger Term Quality

"browser checks" and "UI state" are relevant keywords, but common natural phrases users would say — "browser automation", "Playwright", "Selenium", "end-to-end tests", "UI tests" — are absent. This matches the anchor for some relevant keywords while missing common variations and synonyms, and is below the 4 anchor's good coverage.

3 / 5

Distinctiveness Conflict Risk

"Build reliable browser checks" is somewhat specific to the browser-testing domain but, with no trigger phrases, could still overlap with general test-writing, scraping, or QA skills. It does not reach 4 ('mostly distinct; minor overlap risk') because the technique language alone doesn't clearly separate it from adjacent automation skills.

3 / 5

Total

13

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
sickn33/agentic-awesome-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.