Autonomous data gathering across sources — plans search queries and URLs from a natural-language prompt, executes in parallel within a time budget, optionally extracts structured fields via JSON Schema, and synthesizes results with full step transparency. Use when the user needs data collected from the web with a specific shape, says "gather data", "find pricing for", "collect information about", "extract from multiple sites", or provides a JSON schema for web data.
72
89%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Natural-language data gathering with optional JSON Schema output. Local-first: every fetched page lands in the cache for later reuse.
// Natural language data gathering
{ "prompt": "Find pricing tiers for the top 5 headless CMS platforms" }
// With structured output schema
{
"prompt": "Find pricing for Contentful, Sanity, and Strapi",
"schema": { "type": "object", "properties": { "name": { "type": "string" }, "free_tier": { "type": "string" }, "pro_price": { "type": "string" }, "enterprise": { "type": "string" } } }
}
// With starting URLs
{
"prompt": "Compare features across these CMS platforms",
"urls": ["https://contentful.com/pricing", "https://sanity.io/pricing"],
"max_pages": 6
}| Parameter | Type | Default | When to use |
|---|---|---|---|
prompt | string | required | Natural-language task description |
urls | string[] | none | Seed URLs to include |
schema | object | none | JSON Schema for structured extraction per page |
max_pages | number | 10 | Hard cap on pages fetched (max 100) |
max_time_ms | number | 60000 | Time budget in ms (max 600000) |
stream | boolean | false | Emit progress notifications per step |
max_tokens_out | number | none | Token-budget cap (cl100k-base) |
include_full_markdown | boolean | false | Pages return evidence excerpts by default |
citation_format | string | "numbered" | "numbered" / "json" / "anthropic_tags" |
steps array shows every action with timings.Synthesis follows a fallback ladder: host sampling → optional local language model → deterministic extraction.
Every response includes a steps array:
[
{ "action": "plan", "detail": "Generated 3 search queries", "time_ms": 200 },
{ "action": "search", "detail": "Found 8 results", "time_ms": 5000 },
{ "action": "fetch", "detail": "Fetched 5 pages", "time_ms": 8000 },
{ "action": "extract", "detail": "Extracted schema from 5 sources", "time_ms": 3000 }
]Use steps to debug weak results — if extraction is poor, check which pages were fetched.
research instead.extract instead.max_pages high without time budget — set max_time_ms too.extract.c6ad447
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.