CtrlK
BlogDocsLog inGet started
Tessl Logo

web-scraping-olostep

Web scraping, crawling, and AI-powered answer extraction at scale

45

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/research-tools/capabilities/web-scraping-olostep/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

43%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized but monolithic API dump: it documents every endpoint with parameters and curl examples inline, yet ships two malformed example commands, no async workflow sequencing (polling/validation for crawls and batches), and no progressive disclosure into reference files. It is serviceable for lookup but fails as an operational guide.

Suggestions

Fix the malformed curl examples for Start Crawl and Start Batch — move the body fields inside the -d JSON payload so the commands are executable as written.

Add explicit async workflows with validation checkpoints, e.g. 'Start crawl → poll GET /v1/crawls/{id} until status=completed → fetch /v1/crawls/{id}/pages → retrieve content by retrieve_id', and equivalent for batches.

Move the bulk per-endpoint parameter reference into a references/ file (or lean on the Discover More details endpoint) and keep SKILL.md to setup, the main workflows, and one example per capability.

DimensionReasoningScore

Conciseness

The body is mostly dense parameter lists and curl commands with little concept-explanation padding, but it includes duplication and filler: "Powerful web scraping, crawling, and AI-powered content extraction" restates the description, the Capabilities section repeats the twelve Usage headings verbatim, and long parameter descriptions (e.g., the full default CSS-selector list) could be tightened. This sits at anchor 3 ('mostly efficient but some unnecessary explanation or could be tightened') — above 2 (no heavily padded explanatory prose) but short of 4 due to the redundant capabilities list and boilerplate.

3 / 5

Actionability

Most endpoints have concrete curl commands, but two are syntactically broken: the Start Crawl and Start Batch examples terminate the -d argument at '"path":"/v1/crawls"}' and leave the body fields ("start_url", "max_pages", "items") as orphaned lines outside the quoted JSON, so they fail if copy-pasted. Several endpoints (Batch Items, Crawl Info/Pages, Batch Info, Get Answer, Get Scrape) also show only a '{batch_id}'-style placeholder with no substitution or body/polling detail. This matches anchor 3 ('some concrete guidance but incomplete... missing key details') rather than 4's 'concrete code with minor gaps', since the broken examples are a correctness gap, not just a missing nicety.

3 / 5

Workflow Clarity

The core workflows (start crawl → poll crawl info → fetch pages → retrieve content; start batch → batch info → items → retrieve) are multi-step asynchronous processes, but no sequence is ever laid out — the flow must be inferred from scattered endpoint descriptions — and there is no guidance on polling, checking completion status, or verifying results. Per the guidelines, batch operations without validation steps cap workflow clarity at 3, and this is below that cap: it fits anchor 2 ('rough sequence present but many gaps; validation absent') rather than 3 ('steps listed'), because steps are not listed as a workflow at all.

2 / 5

Progressive Disclosure

The entire ~200-line API reference is inlined in SKILL.md with no references/ bundle and no split into separate files, which matches anchor 3's example ('[200 lines of API reference that could be in a separate file]') — there is clear per-endpoint structure, and the 'Discover More' section offers a live search/details endpoint for deeper info, but the bulk parameter reference clearly belongs in a separate file. Not 2 because the content is well-sectioned and navigable, not a structureless wall; not 4 because no reference files exist and the one-level-deep pattern is absent.

3 / 5

Total

11

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names real capabilities, but it is a bare capability list: it answers 'what' adequately while providing no 'when to use' guidance and only thin natural-keyword coverage. Adding an explicit trigger clause and common synonyms would materially improve it.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to scrape a website, crawl a site for pages/links, or get an AI-synthesized answer from web sources.'

Include natural variations users would actually say — 'scrape a page', 'extract web content', 'get all URLs on a site', 'site map' — to broaden trigger-term coverage.

Drop the marketing filler 'AI-powered' / 'at scale' in favor of a concrete differentiator (e.g., batch scraping, pagination, or structured parsers) to sharpen distinctiveness.

DimensionReasoningScore

Specificity

"Web scraping, crawling, and AI-powered answer extraction" names the domain plus two to three concrete actions (scrape pages, crawl sites, extract answers), matching the anchor 'Names domain and 1-2 concrete actions, but not comprehensive' — it omits operations the skill body documents, such as site maps, batches, and content retrieval. It is above anchor 2 ('Processes PDF files'-level generality) because multiple distinct actions are named, but below anchor 4, which requires a fuller list with only minor gaps.

3 / 5

Completeness

The 'what' is stated clearly (scraping, crawling, AI answer extraction), but there is no 'Use when...' clause or equivalent trigger guidance, which per the judging guidelines caps completeness at 3. Not score 2 because the 'what' half is concrete, not vague; not score 4 because 'when' is entirely absent rather than just imprecise.

3 / 5

Trigger Term Quality

"Web scraping" and "crawling" are natural phrases users would say, but coverage stops there — missing common variations like "scrape a website/page", "data extraction", "spidering", or "get content from a URL", and "AI-powered answer extraction at scale" leans marketing-jargon over user language. This fits anchor 3 ('Some relevant keywords but missing common variations or synonyms') better than anchor 4's 'good keyword coverage'.

3 / 5

Distinctiveness Conflict Risk

"Web scraping, crawling, and AI-powered answer extraction" carves out a recognizable niche (Olostep-style scraping/crawling/answer API) with distinct trigger terms, fitting 'Mostly distinct; minor overlap risk with closely related skills' — it could collide with other web-scraping or fetch-tools skills. Not 5 because the description gives no unique differentiators beyond generic scraping/crawling terms, so overlap with any competing scraper skill remains possible.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 missing, 1 suspicious

Warning

Total

14

/

16

Passed

Repository
gooseworks-ai/goose-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.