robots.txt is a plain-text file at the root of your site that gives crawlers instructions about which URLs they may request. It follows the Robots Exclusion Protocol, standardised as RFC 9309 in 2022. It’s simple, but it’s also one of the easiest ways to accidentally remove a site from search: one Disallow: / left over from a staging site is enough.
What robots.txt can and can’t do
It can:
- stop well-behaved crawlers from fetching certain paths (admin areas, internal search results, endless filter combinations),
- point crawlers to your XML sitemap,
- give different rules to different bots.
It can’t:
- guarantee a page stays out of search results (blocked URLs can still be indexed from links),
- protect private content: the file is public, and bad bots ignore it,
- remove a page that’s already indexed.
Syntax Google supports
Google supports four fields: user-agent, allow, disallow and sitemap. Anything else, including crawl-delay and noindex, is ignored by Googlebot.
User-agent: *
Disallow: /admin/
Disallow: /search
Allow: /admin/public/
Sitemap: https://example.com/sitemap.xml
Key rules:
- Groups: each group starts with one or more
User-agentlines followed by rules. A crawler uses the most specific matching group only, not a combination. If you have aGooglebotgroup, Googlebot ignores the*group. - Paths are case-sensitive and match from the start:
Disallow: /searchalso blocks/search-resultsand/searching. - Wildcards:
*matches any sequence of characters;$anchors the end.Disallow: /*.pdf$blocks all URLs ending in.pdf. - Most specific rule wins: when
AllowandDisallowboth match, Google applies the rule with the longest path; on a tie, the less restrictiveAllowwins. - Empty
Disallow:means “allow everything”. - File size: Google reads only the first 500 KiB.
- Status codes matter: a 404 is treated as “no restrictions”; a prolonged 5xx error can make Google pause crawling the whole site.
Examples
WordPress
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/
Sitemap: https://example.com/sitemap_index.xml
Don’t block /wp-content/ or /wp-includes/: Google needs CSS and JavaScript to render pages properly, and blocking them can hurt how your pages are understood.
Online shop
Faceted navigation (colour, size, price, sort order) can produce millions of near-duplicate URLs and waste crawl capacity. Block the parameters that don’t create unique, valuable pages:
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account/
Disallow: /*?*sort=
Disallow: /*?*price=
Disallow: /*?*sessionid=
Sitemap: https://shop.example.com/sitemap.xml
Keep category pages and filtered pages you actually want ranked (e.g. “red running shoes” if it has its own landing page) crawlable. Build a draft with our robots.txt generator, then test specific URLs against it with the robots.txt tester before publishing.
Blocking AI crawlers
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Google-Extended controls whether your content is used for Google’s AI models; it does not affect normal Google Search crawling or rankings. Check each AI company’s current documentation for the exact user-agent names, as they change over time.
Robots.txt vs noindex
| Goal | Use |
|---|---|
| Save crawl capacity on useless URLs | robots.txt Disallow |
| Keep a page out of search results | <meta name="robots" content="noindex"> or X-Robots-Tag: noindex header |
| Remove a page that’s already indexed | noindex (and keep it crawlable until it drops out) or Search Console’s Removals tool for urgent cases |
| Hide private content | Password protection / authentication |
The classic mistake is combining the two: if you disallow a page in robots.txt, Google can’t crawl it and therefore never sees the noindex tag. The page may stay in the index as a bare URL. Google stopped honouring unofficial noindex rules inside robots.txt in 2019.
Staging and development sites
Staging environments are where robots.txt mistakes start. Teams add Disallow: / to stop a test site being indexed, then the file is copied to production at launch, and organic traffic falls off a cliff a few days later. Two better habits:
- Protect staging with a password (HTTP authentication) or an IP allow-list. That keeps out search engines and people, and nothing needs to change at launch.
- Make robots.txt environment-aware: generate it at build or deploy time so production always gets the production version.
If you do launch with a blocking file by mistake, fix it, then request indexing of key pages in Search Console. Recovery usually starts within days once Google rereads the file.
Subdomains and multiple hosts
Each host needs its own file. blog.example.com/robots.txt doesn’t apply to www.example.com, and the HTTP and HTTPS versions are technically separate too, though after a proper HTTPS redirect only the HTTPS file matters.
Common mistakes
Disallow: /copied from staging to production after launch.- Blocking CSS and JS folders, so Google renders a broken page.
- Using robots.txt to hide sensitive URLs: the file itself advertises them.
- Typos in paths or forgetting that paths are case-sensitive.
- Expecting instant effect: Google usually caches robots.txt for up to 24 hours.
- Serving robots.txt with a 5xx error during maintenance, which can halt crawling.
Checklist
- File is at
https://yourdomain/robots.txt, returns 200 and is under 500 KiB. - No accidental
Disallow: /forUser-agent: *. - CSS, JavaScript and image folders are crawlable.
- Low-value parameter URLs are blocked; valuable landing pages are not.
Sitemap:line points to the correct absolute URL. See our XML sitemap guide.- Pages you want out of search use
noindex, notDisallow. - Important URLs tested in a robots.txt tester after every change.