Robots.txt and sitemap URL checker

Enter a site origin. After you accept the proxy notice, the tool fetches /robots.txt, parses User-agent / Allow / Disallow / Sitemap lines, then fetches the first sitemap (or /sitemap.xml).

This is not Googlebot, Bingbot, or Search Console. Public CORS proxies see the URL. Sites that block proxies will look “down” even when the file exists. We do not spider every loc.

How to Use

  1. Paste a homepage or origin (https://example.com).
  2. Consent to the CORS-proxy disclosure.
  3. Read the robots groups and declared sitemap URLs.
  4. Skim the sitemap loc sample. Resubmit the live sitemap in GSC/Bing yourself after deploys.

Key Features

  • Read-only robots.txt parse
  • Sitemap vs sitemapindex detection
  • Consent-gated public CORS proxies (AllOrigins, CorsProxy, CodeTabs)
  • No site-wide crawl from FlashKit

Frequently Asked Questions

Why did the fetch fail?

The origin may block proxies, require a cookie wall, or return a different robots.txt to datacenter IPs. Try curl locally.

Does this submit my sitemap to Google?

No. You still resubmit in Google Search Console and Bing Webmaster Tools after a release.

Is Disallow: / a guarantee Google will stay out?

It is a request. Malicious bots ignore robots.txt. It also does not hide URLs that are already indexed.

Related Tools