Not getting calls from Google? Find out why. See how it works →
Skip to main content

Robots.txt checker

Enter your website's address to fetch its robots.txt, check it for mistakes that can stop Google crawling your site, and see whether Google may crawl the page you entered.

Free, no signup. Enter your homepage, or any page on the site to see whether that page is blocked.

What is robots.txt?

robots.txt is a plain text file at the root of your site (for example https://www.example.co.uk/robots.txt) that tells crawlers which parts of the site they may crawl. Each group names one or more user agents and lists Allow and Disallow rules by path. A typical WordPress site's file looks like this:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.example.co.uk/sitemap.xml

The file controls crawling, not indexing. In Google's words, robots.txt “is not a mechanism for keeping a web page out of Google”: a blocked URL can still appear in search results, without a description, if other pages link to it. To keep a page out of Google, use a noindex tag or a password, and let Google crawl the page so it can see the noindex.

How Google reads your robots.txt

  • Only at the root. The rules apply to that host, protocol and port only, so the file on www.example.co.uk doesn't cover shop.example.co.uk.
  • No file means no limits. A 404, or any 4xx error except 429, means Google assumes there are no crawl restrictions.
  • A server error pauses crawling. If robots.txt returns a 5xx error, Google stops crawling the site for up to 12 hours, then uses its last cached copy for up to 30 days while it keeps retrying. A broken robots.txt can quietly stop Google crawling your site.
  • Size. Google reads the first 500 KiB and ignores the rest.
  • Paths are case-sensitive. Disallow: /Private/ doesn't block /private/.
  • Unsupported lines are ignored. Google reads User-agent, Allow, Disallow and Sitemap. It ignores Crawl-delay, and a Noindex line in robots.txt does nothing.

Source: How Google interprets the robots.txt specification.

Common robots.txt mistakes and how to fix them

  • “Disallow: /” left in place. New and staging sites are often built with User-agent: * and Disallow: / to keep them out of search. If the site goes live unchanged, Google can't crawl any page. Delete the line (an empty Disallow: allows everything), then ask Google to recrawl your key pages with URL Inspection in Search Console.
  • Blocking CSS and JavaScript. If Google can't fetch the files that lay out your pages, it may not see them as visitors do. Don't block theme, plugin or asset folders unless you're sure the pages work without them.
  • Using robots.txt to hide private pages. The file is public, so it tells anyone where your private folders are, and blocked URLs can still be indexed. Use a password or noindex instead.
  • No Sitemap line. It's optional, but a line such as Sitemap: https://www.example.co.uk/sitemap.xml helps search engines find your sitemap. Check your sitemap while you're at it.

Robots.txt: common questions

Does robots.txt stop a page appearing in Google?

No. It stops Google crawling the page, not indexing it. If other pages link to a blocked URL, it can still appear in results without a description. To keep a page out, add a noindex tag and don't block the page in robots.txt (or Google can't see the tag), or put the page behind a password.

What happens if my site has no robots.txt?

Google treats a missing robots.txt (a 404) as permission to crawl everything. You don't need one, but most sites have one to keep crawlers out of admin areas and to point to the sitemap.

Where should robots.txt go?

At the root of each host, for example https://www.example.co.uk/robots.txt. A file in a subfolder is ignored, and each subdomain (such as shop.example.co.uk) needs its own.

What does "Disallow: /" mean?

Under "User-agent: *", it asks every crawler not to crawl anything on the site. It is the right setting for a private staging site and the wrong one for a live site.

How do I check robots.txt in Google Search Console?

For a domain-level property, open Settings and then the robots.txt report. It shows the robots.txt files Google found, when it last crawled them, and any warnings or errors. To see whether one page is blocked, use URL Inspection.

Does Google obey Crawl-delay?

No. Google ignores the Crawl-delay line and sets its own crawl rate, slowing down if your server struggles. Some other crawlers, such as Bing, do read it.