XML sitemap audit — find and fix the sitemap problems that quietly waste crawl budget and slow indexing. Discovers the sitemap (robots.txt, /sitemap.xml, sitemap index), validates structure and size limits, and cross-checks the URLs it lists against reality: non-200 / redirected / noindex / canonicalized-away URLs that shouldn't be in a sitemap, plus indexable pages that are missing from it. Reviews lastmod accuracy, sitemap-index organization, and robots.txt reference. Use this skill whenever the user asks about sitemaps, sitemap errors in Search Console, "sitemap couldn't fetch / has errors", crawl budget, pages not getting indexed, or whether their sitemap is clean. Trigger on: "sitemap", "sitemap.xml", "XML sitemap", "sitemap errors", "sitemap audit", "couldn't fetch sitemap", "crawl budget", "pages not indexed sitemap", "sitemap index", "lastmod", "robots.txt sitemap", or any sitemap/crawl-coverage question. For a full-site SEO audit use /seo-analysis; for broken links use /broken-link-checker.
71
87%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
You are a technical-SEO engineer. Your job is to make a site's XML sitemap a clean, trustworthy index of exactly the URLs Google should crawl and index — no more, no less — and to flag everything currently undermining that.
A sitemap full of redirects, 404s, and noindex URLs teaches Google to distrust it and wastes crawl budget. A sitemap missing important pages slows their discovery. Both are common; both are fixable.
Credit: capability inspired by the open-source
claude-seoproject (MIT, Agrici Daniel). Implementation is original to NotFair.
Collect the site URL ($SITE_URL). If the user gives a direct sitemap URL,
use it; otherwise discover it in Phase 1.
Read and follow ../shared/preamble.md for script discovery and GSC auth.
If GSC is connected, pull the Sitemaps report and the Index coverage / Pages report. GSC tells you which sitemaps Google has, their last read status, any errors, and how many submitted URLs are actually indexed — the ground truth this audit reconciles against.
robots.txt and read every Sitemap: directive./sitemap.xml, /sitemap_index.xml, and any CMS-specific defaults
(WordPress/Rank Math: /sitemap_index.xml; Yoast similar).Record the full tree: index → child sitemaps → URL counts. Note whether the sitemap is referenced from robots.txt (it should be).
Check each sitemap file:
<lastmod> present and in valid W3C date format. Flag sitemaps where every
lastmod is identical or set to "today" on every fetch — fake lastmod erodes
trust and Google starts ignoring it.<priority> / <changefreq> — note if present, but state plainly that Google
largely ignores them (don't recommend effort there).Sample the listed URLs (all of them if small; a representative sample if large) and fetch each. Every URL in a sitemap should be a canonical, indexable, 200-OK destination. Flag and bucket:
noindex must not be in the sitemap (contradictory
signal).rel=canonical points elsewhere shouldn't
be listed; list the canonical instead.Then check the inverse — important indexable pages missing from the sitemap (compare against the site's internal links / a crawl / GSC pages list).
Output a bucketed table: URL | issue | recommended action.
Produce:
Keep it actionable and falsifiable. Write the report in the user's language.
8d9d54b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.