What is robots.txt?
robots.txt is a plain text file at the root of your site (for example https://www.example.co.uk/robots.txt) that tells crawlers which parts of the site they may crawl. Each group names one or more user agents and lists Allow and Disallow rules by path. A typical WordPress site's file looks like this:
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://www.example.co.uk/sitemap.xml
The file controls crawling, not indexing. In Google's words, robots.txt “is not a mechanism for keeping a web page out of Google”: a blocked URL can still appear in search results, without a description, if other pages link to it. To keep a page out of Google, use a noindex tag or a password, and let Google crawl the page so it can see the noindex.
How Google reads your robots.txt
- Only at the root. The rules apply to that host, protocol and port only, so the file on
www.example.co.ukdoesn't covershop.example.co.uk. - No file means no limits. A 404, or any 4xx error except 429, means Google assumes there are no crawl restrictions.
- A server error pauses crawling. If robots.txt returns a 5xx error, Google stops crawling the site for up to 12 hours, then uses its last cached copy for up to 30 days while it keeps retrying. A broken robots.txt can quietly stop Google crawling your site.
- Size. Google reads the first 500 KiB and ignores the rest.
- Paths are case-sensitive.
Disallow: /Private/doesn't block/private/. - Unsupported lines are ignored. Google reads
User-agent,Allow,DisallowandSitemap. It ignoresCrawl-delay, and aNoindexline in robots.txt does nothing.
Common robots.txt mistakes and how to fix them
- “Disallow: /” left in place. New and staging sites are often built with
User-agent: *andDisallow: /to keep them out of search. If the site goes live unchanged, Google can't crawl any page. Delete the line (an emptyDisallow:allows everything), then ask Google to recrawl your key pages with URL Inspection in Search Console. - Blocking CSS and JavaScript. If Google can't fetch the files that lay out your pages, it may not see them as visitors do. Don't block theme, plugin or asset folders unless you're sure the pages work without them.
- Using robots.txt to hide private pages. The file is public, so it tells anyone where your private folders are, and blocked URLs can still be indexed. Use a password or noindex instead.
- No Sitemap line. It's optional, but a line such as
Sitemap: https://www.example.co.uk/sitemap.xmlhelps search engines find your sitemap. Check your sitemap while you're at it.