CtrlK
BlogDocsLog inGet started
Tessl Logo

web-scraping

Scrape websites, extract structured data, and automate browsers. Use when asked to scrape, extract, crawl, parse, or pull data from web pages or any URL.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/research-tools/capabilities/web-scraping/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a broad, well-organized API catalog with concrete commands and clear per-task routing ('Best for', Tips), but it is not copy-paste reliable: multiple curl examples have JSON bodies placed outside the -d argument, and the closing 'Discover More' section is garbled. Combined with ~40-fold duplication of the auth boilerplate, zero error-handling guidance for async polls and batch jobs, and no use of reference files for the five API catalogs, the skill is serviceable but needs tightening and verification steps.

Suggestions

Fix the malformed curl examples: in the smartscraper, crawl, riveter /v1/run, and batch blocks, the JSON body appears outside the -d '...' argument (the quote closes at "path":"/v1/..." and body keys dangle unquoted); rewrite each as a single complete -d '{...}' payload, and clean up the 'Discover More' section where stray text ('api show olostep') and a stray backtick break the commands.

State the shared curl invocation once (base URL, Bearer auth header, api/path/body structure) and show only the per-endpoint JSON payloads in the examples — this removes the four-line header block repeated ~40 times and roughly halves the file's token cost.

Add validation and error-recovery guidance for the async and batch workflows: how to interpret poll responses (pending vs. completed vs. failed), what to do on error, and a checkpoint before retrieving batch/crawl results.

Move the full endpoint catalogs for each API (Scrapegraph, Olostep, Riveter, Brand.dev, Notte) into per-API reference files under references/, keeping SKILL.md as an overview with 'Best for' routing and one or two key examples per API.

DimensionReasoningScore

Conciseness

The identical four-line curl auth-header block (URL, Authorization, Content-Type, api/path prefix) is repeated in nearly every one of ~40 examples, inflating the file well past what the content earns. It avoids explaining concepts Claude already knows, but the boilerplate duplication means it 'could be tightened' — factoring the shared invocation into one stated pattern would cut the file roughly in half.

3 / 5

Actionability

Commands are concrete curl calls, but a significant number are malformed and not executable as written: several blocks (e.g., the smartscraper, crawl, and riveter /v1/run examples) close the -d quote before the JSON body, leaving body keys like "website_url" outside the request; the 'Discover More' section appends stray text ('api show olostep') after the -d argument and contains a stray backtick. This is 'concrete guidance but incomplete; missing key details' rather than mostly-executable with minor gaps.

3 / 5

Workflow Clarity

Multi-step flows are clearly labeled (Notte 'Step 1–5', Olostep crawl start/status/pages/retrieve), giving a real sequence. However, validation is absent throughout: no guidance on checking poll responses for errors or completion, no error-recovery loop, and the batch-scraping operations (Olostep batches, async crawls) run with no verification steps — batch operations without validation cap this dimension at 3 per the rubric.

3 / 5

Progressive Disclosure

The file has good section structure (numbered per-API sections with 'Best for' lines and a Tips section), but there are no bundle files at all: full endpoint catalogs for five distinct APIs (~440 lines) are inlined in SKILL.md, where the rubric expects that bulk to live one level deep in reference files. This fits 'some structure but content that should be separate is inline'.

3 / 5

Total

12

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete actions, and an explicit 'Use when' clause with natural trigger verbs. The only weaknesses are modest — slightly incomplete action coverage relative to the skill's full capability set and a few missing colloquial trigger phrasings.

DimensionReasoningScore

Specificity

"Scrape websites, extract structured data, and automate browsers" names three concrete actions, matching the 'several specific actions; minor gaps' anchor. Not a 5 because coverage has gaps — the body also covers brand assets, markdown conversion, and batch scraping, none of which the description mentions.

4 / 5

Completeness

It explicitly answers both what ("Scrape websites, extract structured data, and automate browsers") and when ("Use when asked to scrape, extract, crawl, parse, or pull data from web pages or any URL") with concrete trigger phrases, matching the top anchor. Not lower: both halves are present and explicit, not implied.

5 / 5

Trigger Term Quality

"scrape, extract, crawl, parse, or pull data from web pages or any URL" gives good natural-verb coverage with synonyms, fitting the 'good keyword coverage; a few natural terms missing' anchor. It falls short of comprehensive — common phrasings like 'spider', 'harvest', or 'get the content off this page' are absent.

4 / 5

Distinctiveness Conflict Risk

Web scraping is a clear niche with distinct triggers ("web pages", "any URL"), matching 'mostly distinct; minor overlap risk'. Not a 5 because generic verbs like 'extract' and 'parse' could also fire for document- or data-extraction skills.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
gooseworks-ai/goose-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.