Use this skill whenever a user wants to crawl one or more web pages or a whole site and turn them into an Algolia index using the Algolia CLI — especially for RAG, AI search, semantic search, or Agent Studio retrieval. Triggers: "index my website/docs with Algolia", "set up the Algolia Crawler", "crawl this page for RAG", "scrape my site into Algolia", "build a knowledge base for my AI agent from these URLs", writing or debugging a crawler recordExtractor, or handling JavaScript-rendered pages that won't index. It guides ingestion end-to-end with `algolia crawler` commands: inspect the page, write a RAG-optimized recordExtractor, validate with `algolia crawler test` BEFORE indexing, apply index settings explicitly, then reindex. Do NOT use for building the chatbot/agent layer itself (use algobot-cli), raw record/synonym/settings ops on an existing index (use algolia-cli), frontend search UI (use instantsearch), or read-only search/analytics (use algolia-mcp).
80
100%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Turn web pages into a RAG-optimized Algolia index with the Algolia Crawler, driven entirely by the Algolia CLI (algolia crawler …). The Crawler visits URLs, extracts content with a JavaScript recordExtractor, and writes Algolia records on a schedule — no scraping code to maintain.
The single most important idea: crawling for RAG is not the same as crawling for site search. A RAG record is a chunk of text that will be dropped into an LLM's context window in isolation. That changes how you shape every record (see The RAG record mental model).
| Need | Skill |
|---|---|
| Crawl/ingest web pages into an index (this workflow) | algolia-crawler |
| Build the chatbot / agent / RAG app on top of the index | algobot-cli |
| Raw record/settings/synonym/rule ops on an existing index | algolia-cli |
| Read-only search, analytics, recommendations | algolia-mcp |
| Frontend search UI (autocomplete, results, facets) | instantsearch |
Rule of thumb: if the user wants to get web content into Algolia, you're in the right skill. Once the index exists and they want to ask questions over it, hand off to algobot-cli.
Each record is retrieved and shown to an LLM alone, without the surrounding page. So optimize for that:
content, even for structured data. This is the highest-leverage move. Don't store only {model: "X", score: 0.87} — also synthesize a sentence: "Model X scores 0.87 on relevance…". Both keyword and vector retrieval, and the LLM itself, want prose.content, keep clean fields (category, type, date, numeric values) so retrieval can be scoped with filters and facets.If you internalize only one thing from this skill, make it this list. Everything else is mechanics.
Follow these steps in order. Full commands, config, and code live in the references — read them as you reach each step.
recordExtractor that emits RAG records. For docs and prose, the recommended default is Algolia's Markdown helpers (helpers.markdown + helpers.splitTextIntoRecords) — they preserve headings/lists/code and are Algolia's own AskAI/RAG pattern. Hand-roll an extractor when you need per-entity structured records (tables, catalogs). Either way, follow the mental model above. See record-extractor.md.algolia crawler test BEFORE indexing. It runs your config against a live URL and returns the records it would create, without writing anything. Use it as your feedback loop — and as a DOM inspector to fix selectors against the real rendered markup. See workflow.md.initialIndexSettings to configure the index — in practice it frequently does not apply. Set searchableAttributes, attributesForFaceting, etc. yourself with algolia settings import after the first crawl. See rag-index-settings.md.algolia search) for a few realistic RAG questions before trusting it.Crawling a whole site or section (one URL → all its sub-pages)? The crawler discovers sub-pages via startUrls/sitemaps/discoveryPatterns — but the config must be shaped from all the page types it will hit, not just the start page. A config fitted to the landing page produces junk on the article/reference/blog templates it never saw. Discover the URL set, group it into templates, sample a representative page per template, and validate the extractor against each. See site-crawls.md.
The complete, copy-adaptable end-to-end run (create → test → settings → reindex → verify) is in workflow.md.
These are the non-obvious failures. Both references cover them, but they're worth stating up front:
Loading… placeholders in its initial HTML, or the numbers/table appear only after load, you must enable renderJavaScript. A crawl without it indexes nothing useful. (javascript-rendering.md)data-tip="Score: 79.8%") while the visible cell shows only a bar or a delta. Extract from the attribute, not the visible text. (record-extractor.md)Everything runs through algolia crawler … (and algolia settings / algolia search for the index side). Full cheatsheet and the gotchas below are in cli.md.
algolia crawler create <name> -F config.json # create (prints nothing on success)
algolia crawler test <id> --url <url> [-F cfg] # extract records WITHOUT indexing
algolia crawler reindex <id> # start a crawl that writes records
algolia crawler get <id> # inspect crawler + config/status
algolia crawler stats <id> # crawl status summaryTwo CLI realities to plan around (details in cli.md):
renderJavaScript: true (boolean). The CLI only accepts the boolean form; the object form ({ enabled, patterns, waitTime }) makes get/list/create -F fail to parse. Boolean true uses the default render wait, which is enough for most pages — confirm with crawler test (empty values mean the page needs more render time).create prints nothing and there's no config-update or delete command. After create, recover the id from algolia crawler list (get needs a UUID — it doesn't accept a name). If list errors because another crawler in the account uses a non-boolean renderJavaScript, there's no pure-CLI recovery — look the id up via the Crawler REST list endpoint (GET /1/crawlers?name=<name>). To change a config, re-create; to delete, use the dashboard or REST DELETE.algolia crawler command cheatsheet, and the CLI gotchas in detail.0953d1b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.