CtrlK
BlogDocsLog inGet started
Tessl Logo

stagehand-automation

AI-powered browser automation using Stagehand v3 and Claude. Use when building self-healing tests, AI agents, dynamic web automation, or when traditional selectors break frequently due to UI changes.

59

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/stagehand-automation/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable body with a clean quick-start and concrete code throughout, but it is padded with marketing language and redundant self-healing explanations, and its progressive disclosure is broken: every referenced file is missing while both actual bundle files go unreferenced, with reference-worthy detail inlined instead.

Suggestions

Fix the References section to point at the files that actually exist (`references/api-reference.md`, `references/troubleshooting.md`) instead of the three nonexistent paths.

Move the detailed API examples, model-selection/cost tables, and MCP server setup into the existing reference files, keeping SKILL.md as a lean overview with the quick start and pointers.

Trim redundant sections (When Self-Healing Activates, Caching, Traditional vs Stagehand) and drop marketing claims like 'state-of-the-art' and '44% faster than v2' that add tokens without actionable value.

DimensionReasoningScore

Conciseness

The body is code-heavy and much of it earns its place, but it carries unnecessary material: marketing claims ('state-of-the-art', '44% faster than v2', 'Key Innovation'), the self-healing concept re-explained across four sections (Overview, Self-Healing Patterns, When Self-Healing Activates, Caching), and a volatile 'Estimated Costs' table. This fits 'Mostly efficient but includes some unnecessary explanation or could be tightened' rather than the minor-instances 4.

3 / 5

Actionability

The Quick Start is a complete, runnable script with install and run commands, and the act/extract/observe sections give concrete, mostly copy-paste-ready examples. Minor gaps keep it below 5: the hybrid example uses `z` without importing zod, the cost-optimization example references an undefined `complexSchema`, and the install step omits ts-node/typescript needed for `npx ts-node`.

4 / 5

Workflow Clarity

The Quick Start is a clearly sequenced 1-4 flow (install, configure, write, run), and the Error Handling section supplies timeout handling and a retry feedback loop. It is not a destructive/batch operation so the validation cap does not apply, but the main workflow lacks explicit verification checkpoints (e.g., confirm the browser launched or the action succeeded before proceeding), matching 'Clear sequence with most checkpoints present; minor validation gaps'.

4 / 5

Progressive Disclosure

Scored against the actual bundle: the References section points to three files (`references/stagehand-v3-guide.md`, `references/claude-integration.md`, `references/self-healing-patterns.md`) that do not exist, while the two real bundle files (`references/api-reference.md`, `references/troubleshooting.md`) are never mentioned, so navigation is broken in both directions. Compounding this, ~200 lines of API detail, model-selection, cost tables, and MCP setup are inlined in the 420-line body where they clearly belong in the reference files, matching 'Minimal structure; content that clearly belongs in separate files is inlined'.

2 / 5

Total

13

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' clause and natural trigger phrases tied to a distinct niche. Its main weakness is specificity: it describes the domain and use cases rather than the concrete actions (act/extract/observe) the skill performs.

Suggestions

Replace the generic 'AI-powered browser automation' capability statement with 1-2 concrete actions, e.g. 'Drive browsers with natural-language actions, extract structured page data with Zod schemas, and observe page state'.

Add common synonym triggers users actually say, such as 'web scraping', 'E2E tests', or 'end-to-end tests', to broaden natural trigger coverage.

Qualify the broad 'AI agents' trigger (e.g. 'browser-driving AI agents') to reduce overlap with general agent-building skills.

DimensionReasoningScore

Specificity

The description names the domain ("AI-powered browser automation using Stagehand v3 and Claude") and gestures at capabilities via use cases ("building self-healing tests", "dynamic web automation"), but never states concrete actions like acting on elements, extracting structured data, or observing page state. It matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'; a 4 would require several specific actions listed.

3 / 5

Completeness

It explicitly answers both questions: the 'what' ("AI-powered browser automation using Stagehand v3 and Claude") and an explicit 'when' with concrete trigger phrases ("Use when building self-healing tests, AI agents, dynamic web automation, or when traditional selectors break frequently due to UI changes"). This mirrors the 5-anchor example structure of a clear capability statement followed by an explicit 'Use when...' clause.

5 / 5

Trigger Term Quality

Natural trigger phrases are present: "browser automation", "self-healing tests", "AI agents", "dynamic web automation", "when traditional selectors break frequently due to UI changes". Coverage is good but a few natural variations users would say are missing (e.g. "web scraping", "end-to-end/E2E tests", "Playwright"), so it fits 'Good keyword coverage; a few natural terms missing' rather than the comprehensive 5.

4 / 5

Distinctiveness Conflict Risk

The Stagehand/self-healing niche is fairly distinct from generic web-dev or testing skills. However, the trigger "AI agents" is broad and could overlap with non-browser agent-building skills, and 'browser automation' overlaps with generic Playwright/Selenium skills, so it fits 'Mostly distinct; minor overlap risk' rather than the minimal-conflict 5.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.