A sitemap should list the pages you want in search, at the address each one lives at, and nothing else. Every wrong entry costs a crawl and teaches search engines to trust your sitemap less. Most sitemaps have some: URLs that redirect or 404, pages marked noindex, pages whose canonical is another URL, and lastmod dates that are all the same.

Enter your site and this tool reads robots.txt and every sitemap it names (indexes and gzip files too), checks the URLs, and tells you exactly what to change: “Remove these 37 URLs from /post-sitemap.xml”, “Replace these with the URL they redirect to”, “Fix the canonical on this page”.

Sign in to run it. It’s free for up to 100 URLs. Bigger runs cost 1 credit ($1) per 1,000 URLs checked, and you only pay for the URLs it checks: if your site lists 300 URLs and you allow 5,000, you pay 1 credit and the rest comes back. A run that fails gives all its credits back.

Loading

This did not load. Reload the page to try again.

Sign in free to run Sitemap health on your own site

Free for small sites. No card needed.

We’ll email you when it’s ready. Change

Your recent runs

SiteResultDate

You can close this page. The run keeps going, and the report stays in your reports.

Sitemaps
Listed
Checked
Problems
Elapsed
  1. robots.txt
  2. Sitemaps
  3. Check URLs
  4. Home page links
  5. Report

Starting

Checks up to 10,000 URLs. The crawler obeys robots.txt and Crawl-delay. Each report is kept in your account, where only you can open it, until you delete the account.

What it checks

Every sitemap file:

  • Found: robots.txt names it, or it is at a usual address such as /sitemap.xml. If robots.txt does not name it, you are told to add the line.
  • Loads with a 200, at its own address (no redirect), served as XML.
  • Is a real sitemap or sitemap index, with no more than 50,000 URLs and no more than 50 MB.
  • A sitemap index does not list another index, and gives each sitemap a lastmod.

Every URL listed:

  • Listed once, on your own host.
  • Has a lastmod in the right format that is not in the future, and not the same date as nearly every other URL (a sign the date is when the sitemap was built, not when the page changed).

Up to your limit of URLs, taken from every sitemap in turn so each section is covered:

  • Not blocked by robots.txt.
  • Answers 200 without a redirect. A redirect’s final URL is shown, so you can swap it in.
  • Not marked noindex, in a robots meta tag or the X-Robots-Tag header.
  • Its canonical URL is itself. A trailing slash or http counts as a different URL.

And the other way round: pages your home page links to that load, can be indexed, and are in no sitemap.

What you get

  • One card per fix, most urgent first, with the exact change to make and every URL it covers.
  • A table of your sitemaps with how many URLs each lists, how many were checked and how many need a fix.
  • CSV and JSON downloads, with one row per URL and its fix.
  • “Copy for agent” gives your AI agent a short index and one fix at a time, so it can change the code or plugin that builds your sitemap.
  • An email when the report is ready. You can turn that off in your account settings.

How the crawl works

  • The site is crawled by ianbot. It obeys robots.txt and Crawl-delay. If your firewall blocks ianbot, the ianbot page shows how to allow it.
  • It reads only the <head> of each page, which is where noindex and the canonical live.
  • It reads up to 200 sitemap files and 250,000 listed URLs per run.

Pricing and privacy

  • Free for up to 100 URLs. Above that, 1 credit ($1) per 1,000 URLs checked, rounded up. Credits are charged when the run starts; what it did not need comes back when it ends, and everything comes back if it fails.
  • Each run gets its own link, which only you can open. The report is kept in your reports until you delete your account. Your recent runs are listed above the form.

Newsletter

New posts, daily by default. Change any time.

Join 9,363 active subscribers