webtrajans
en

Robots.txt: syntax, examples and the mistakes to avoid

Robots.txt tells crawlers which parts of your site not to crawl. A single wrong line can hide your whole site from Google, so it pays to understand the rules.

Updated: 5 min read

robots.txt is a plain-text file at the root of your site that gives crawlers instructions about which URLs they may request. It follows the Robots Exclusion Protocol, standardised as RFC 9309 in 2022. It’s simple, but it’s also one of the easiest ways to accidentally remove a site from search: one Disallow: / left over from a staging site is enough.

What robots.txt can and can’t do

It can:

  • stop well-behaved crawlers from fetching certain paths (admin areas, internal search results, endless filter combinations),
  • point crawlers to your XML sitemap,
  • give different rules to different bots.

It can’t:

  • guarantee a page stays out of search results (blocked URLs can still be indexed from links),
  • protect private content: the file is public, and bad bots ignore it,
  • remove a page that’s already indexed.

Syntax Google supports

Google supports four fields: user-agent, allow, disallow and sitemap. Anything else, including crawl-delay and noindex, is ignored by Googlebot.

User-agent: *
Disallow: /admin/
Disallow: /search
Allow: /admin/public/

Sitemap: https://example.com/sitemap.xml

Key rules:

  • Groups: each group starts with one or more User-agent lines followed by rules. A crawler uses the most specific matching group only, not a combination. If you have a Googlebot group, Googlebot ignores the * group.
  • Paths are case-sensitive and match from the start: Disallow: /search also blocks /search-results and /searching.
  • Wildcards: * matches any sequence of characters; $ anchors the end. Disallow: /*.pdf$ blocks all URLs ending in .pdf.
  • Most specific rule wins: when Allow and Disallow both match, Google applies the rule with the longest path; on a tie, the less restrictive Allow wins.
  • Empty Disallow: means “allow everything”.
  • File size: Google reads only the first 500 KiB.
  • Status codes matter: a 404 is treated as “no restrictions”; a prolonged 5xx error can make Google pause crawling the whole site.

Examples

WordPress

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/

Sitemap: https://example.com/sitemap_index.xml

Don’t block /wp-content/ or /wp-includes/: Google needs CSS and JavaScript to render pages properly, and blocking them can hurt how your pages are understood.

Online shop

Faceted navigation (colour, size, price, sort order) can produce millions of near-duplicate URLs and waste crawl capacity. Block the parameters that don’t create unique, valuable pages:

User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account/
Disallow: /*?*sort=
Disallow: /*?*price=
Disallow: /*?*sessionid=

Sitemap: https://shop.example.com/sitemap.xml

Keep category pages and filtered pages you actually want ranked (e.g. “red running shoes” if it has its own landing page) crawlable. Build a draft with our robots.txt generator, then test specific URLs against it with the robots.txt tester before publishing.

Blocking AI crawlers

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Google-Extended controls whether your content is used for Google’s AI models; it does not affect normal Google Search crawling or rankings. Check each AI company’s current documentation for the exact user-agent names, as they change over time.

Robots.txt vs noindex

Goal Use
Save crawl capacity on useless URLs robots.txt Disallow
Keep a page out of search results <meta name="robots" content="noindex"> or X-Robots-Tag: noindex header
Remove a page that’s already indexed noindex (and keep it crawlable until it drops out) or Search Console’s Removals tool for urgent cases
Hide private content Password protection / authentication

The classic mistake is combining the two: if you disallow a page in robots.txt, Google can’t crawl it and therefore never sees the noindex tag. The page may stay in the index as a bare URL. Google stopped honouring unofficial noindex rules inside robots.txt in 2019.

Staging and development sites

Staging environments are where robots.txt mistakes start. Teams add Disallow: / to stop a test site being indexed, then the file is copied to production at launch, and organic traffic falls off a cliff a few days later. Two better habits:

  • Protect staging with a password (HTTP authentication) or an IP allow-list. That keeps out search engines and people, and nothing needs to change at launch.
  • Make robots.txt environment-aware: generate it at build or deploy time so production always gets the production version.

If you do launch with a blocking file by mistake, fix it, then request indexing of key pages in Search Console. Recovery usually starts within days once Google rereads the file.

Subdomains and multiple hosts

Each host needs its own file. blog.example.com/robots.txt doesn’t apply to www.example.com, and the HTTP and HTTPS versions are technically separate too, though after a proper HTTPS redirect only the HTTPS file matters.

Common mistakes

  • Disallow: / copied from staging to production after launch.
  • Blocking CSS and JS folders, so Google renders a broken page.
  • Using robots.txt to hide sensitive URLs: the file itself advertises them.
  • Typos in paths or forgetting that paths are case-sensitive.
  • Expecting instant effect: Google usually caches robots.txt for up to 24 hours.
  • Serving robots.txt with a 5xx error during maintenance, which can halt crawling.

Checklist

  1. File is at https://yourdomain/robots.txt, returns 200 and is under 500 KiB.
  2. No accidental Disallow: / for User-agent: *.
  3. CSS, JavaScript and image folders are crawlable.
  4. Low-value parameter URLs are blocked; valuable landing pages are not.
  5. Sitemap: line points to the correct absolute URL. See our XML sitemap guide.
  6. Pages you want out of search use noindex, not Disallow.
  7. Important URLs tested in a robots.txt tester after every change.

Frequently asked questions

Does robots.txt stop a page from appearing in Google?

Not reliably. Disallow stops crawling, but Google can still index a blocked URL without its content if other pages link to it. To keep a page out of search results, allow crawling and use a noindex meta tag or X-Robots-Tag header.

Where must the robots.txt file be located?

At the root of each host and protocol, e.g. https://example.com/robots.txt. A file at https://example.com/folder/robots.txt is ignored, and subdomains like shop.example.com need their own file.

Does Google support crawl-delay?

No. Googlebot ignores crawl-delay; Bing and Yandex do honour it. Google adjusts its crawl rate automatically based on how your server responds.

Can I block AI crawlers with robots.txt?

You can disallow user agents that announce themselves, such as GPTBot, ClaudeBot, CCBot or Google-Extended. Compliance is voluntary, so it works for well-behaved crawlers but is not access control.

Related guides