CtrlK
BlogDocsLog inGet started
Tessl Logo

stagehand-automation

AI-powered browser automation using Stagehand v3 and Claude. Use when building self-healing tests, AI agents, dynamic web automation, or when traditional selectors break frequently due to UI changes.

54

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/stagehand-automation/SKILL.md

The canonical home for this skill is stagehand-automation in fernandezbaptiste/Skrillz

SKILL.md
Quality
Evals
Security

Quality

Content

52%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is strong on executable guidance — a runnable quick start, concrete per-API examples, and real error-handling/retry patterns — but it underuses its bundle: every reference pointer is broken, the two actual reference files are orphaned, and large reference-grade content (full API examples, cost tables, MCP setup) is inlined, with marketing claims and dated model IDs adding token cost without value.

Suggestions

Fix the References section to point at the files that actually exist — `references/api-reference.md` and `references/troubleshooting.md` — instead of the nonexistent `stagehand-v3-guide.md`, `claude-integration.md`, and `self-healing-patterns.md`.

Move the full Core API example sets, the Estimated Costs table, MCP integration, and model-selection detail into the reference files, keeping SKILL.md to a lean quick start plus a one-line pointer per advanced topic.

Cut the duplicated Core APIs listing from the Overview, the repeated model config in 'Claude Integration', and the time-sensitive marketing claims ('44% faster than v2', dated model IDs) — or move version/model specifics into a clearly-labeled 'deprecated/old patterns' style section.

DimensionReasoningScore

Conciseness

The 425-line body has several padded sections: the Core APIs are listed twice (once in the Overview's 'Core APIs' bullet list, then again in full), the Quick Start's model config is repeated verbatim in 'Claude Integration', and there are marketing/explanatory sections Claude does not need ('state-of-the-art... 44% faster than v2', 'integrates seamlessly', the 'Traditional vs Stagehand' and 'When Self-Healing Activates' explainers). Time-sensitive specifics — dated model IDs like 'claude-sonnet-4-20250514' and 'claude-3-5-haiku-20241022' plus an 'Estimated Costs' table — appear outside any 'old patterns'/'deprecated' section. This matches 'noticeably verbose; several unnecessary explanations or padded sections'; it is above 1 because the bulk is code rather than concept tutorials, and below 3 because the padding and duplication go beyond a single stray explanation.

2 / 5

Actionability

The Quick Start is copy-paste ready (npm install, .env key, a complete runnable script, `npx ts-node`), and the act()/extract()/observe() sections, the error-handling section with a working `actWithRetry` loop, and the hybrid Playwright example are all real, executable TypeScript. It falls short of 5 on minor gaps: the hybrid example uses `z.object` without importing zod in that snippet, and the cost-optimization example references an undefined `complexSchema`. It is well above 3 because nothing is pseudocode and the common cases are covered with concrete code.

4 / 5

Workflow Clarity

The Quick Start gives a clear numbered sequence — '1. Install Stagehand', '2. Configure Claude API', '3. Write First Automation', '4. Run' — each with its command, and the Error Handling section supplies a genuine feedback loop (try/catch on timeout/multiple-match errors plus a retry pattern with backoff). It misses 5 because the happy-path workflow has no explicit validation checkpoint (e.g., verifying the extract() result or asserting the action succeeded before proceeding); it is above 3 because sequence and error-recovery are explicitly present, not implicit.

4 / 5

Progressive Disclosure

Scored against the actual bundle: the References section cites `references/stagehand-v3-guide.md`, `references/claude-integration.md`, and `references/self-healing-patterns.md` — none of which exist — while the real files present (`references/api-reference.md`, `references/troubleshooting.md`) are never mentioned, leaving the troubleshooting guide undiscoverable. Meanwhile ~200 lines of material that clearly belongs in those reference files (full API example sets, the cost table, MCP integration, model-selection detail) are inlined in SKILL.md. This matches 'content that clearly belongs in separate files is inlined; or references are buried' — broken pointers plus orphaned real files is worse than the 3-anchor's 'references present but not clearly signaled'. It is not 1 because the body itself is well-sectioned with clear headers and is navigable as a standalone document.

2 / 5

Total

12

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description with an explicit and well-crafted 'Use when...' trigger clause (the selectors-breaking scenario is especially natural). The main weakness is that the 'what' stays at the level of use-case framing rather than naming the concrete operations the skill performs.

DimensionReasoningScore

Specificity

The description names the domain ('AI-powered browser automation using Stagehand v3 and Claude') and gestures at capabilities via use cases ('self-healing tests, AI agents, dynamic web automation'), but never names a concrete action — there is no equivalent of 'extract data' or 'fill forms', and the core act/extract/observe operations are absent. It sits at the 'names domain and 1-2 concrete actions, but not comprehensive' anchor; it is below 4 because the capability list is use-case framing rather than specific actions, and above 2 because the domain and intent are clearly stated, not generic.

3 / 5

Completeness

Both parts are explicit: the 'what' is 'AI-powered browser automation using Stagehand v3 and Claude' and the 'when' is a full clause — 'Use when building self-healing tests, AI agents, dynamic web automation, or when traditional selectors break frequently due to UI changes'. It misses 5 because the 'what' is a single generic capability sentence without concrete actions (compare the 5-anchor's 'Extract text and tables, fill forms, merge documents'), while the 'when' is as explicit as the 4-anchor example ('Use when working with PDF files'). Not 3, since the 'when' is present and explicit, not weakly implied.

4 / 5

Trigger Term Quality

Natural phrases a user would actually say are present: 'browser automation', 'self-healing tests', 'dynamic web automation', and the excellent scenario trigger 'when traditional selectors break frequently due to UI changes'. It falls short of 5 because common synonyms are missing — no 'web scraping', 'E2E tests', 'Playwright'/'Selenium' (the tools users migrating from would name), or 'form filling'. It is clearly above 3 ('Works with PDF files'-level coverage) since several distinct natural terms are covered.

4 / 5

Distinctiveness Conflict Risk

Naming the specific framework ('Stagehand v3') carves out a clear niche with distinct triggers ('self-healing tests', 'selectors break'), so conflict risk is low — matching 'mostly distinct; minor overlap risk'. It is not 5 because 'AI agents' and 'dynamic web automation' are broad and could pull this skill in for general agent-building or generic scraping/automation requests where a Playwright or scraping skill would be the better match.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

15

/

16

Passed

Repository
fernandezbaptiste/Skrillz
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.