The robots.txt file is a plain text file at the root of your domain (e.g., yourdomain.com/robots.txt) that tells search engine crawlers which parts of your site they’re allowed — or not allowed — to access.
Basic syntax: Each rule specifies a User-agent (which crawler it applies to) and Disallow or Allow directives (which paths to block or permit). For example, blocking a staging folder from all crawlers looks like:
User-agent: *
Disallow: /staging/
Common mistakes: The most damaging error is accidentally blocking your entire site with Disallow: / after a site migration or redesign — this tells every crawler to stay away from everything. Another frequent issue is blocking CSS or JavaScript files that Google needs to render the page properly, which can hurt how your content is understood.
What robots.txt does NOT do: It doesn’t remove pages from Google’s index — a blocked page can still appear in search results (usually without a description) if other sites link to it. To fully keep a page out of search results, use a noindex meta tag instead, on a page that robots.txt allows crawlers to reach.
Best practices: Keep the file simple, always include your sitemap URL at the bottom, check the robots.txt report in Search Console after deploying, and audit it after every major site migration.
A practical robots.txt example
Here is a typical file for a small business site or online store:
User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /checkout/
Disallow: /*?sessionid=
Sitemap: https://example.com/sitemap.xml
This keeps crawlers out of private and low-value areas, blocks session-ID URLs that create endless duplicates, and points every crawler to the sitemap.
Wildcards and pattern matching
Google and Bing support two special characters:
*matches any sequence of characters.Disallow: /*?filter=blocks every URL containing?filter=.$marks the end of a URL.Disallow: /*.pdf$blocks URLs that end in .pdf, but not a page like/pdf-guide/.
When an Allow and a Disallow rule both match a URL, Google follows the most specific (longest) rule. This lets you block a folder but allow one file inside it.
Controlling AI crawlers
Robots.txt is also how you decide whether AI crawlers can use your content. Common user-agents include GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot and Google-Extended, which controls whether Google may use your content for Gemini training. Google-Extended does not affect Google Search or AI Overviews; those use the normal Googlebot. Before blocking AI crawlers, decide whether being cited in AI answers matters to your business. For most brands that want visibility in ChatGPT and Perplexity, blocking them works against that goal.
How to test robots.txt safely
Google retired the old robots.txt Tester in 2023. Today you can:
- Open the robots.txt report in Search Console (Settings → robots.txt) to see the version Google last fetched and any parsing errors.
- Use URL Inspection on an important page to confirm it is not blocked.
- Test new rules in a third-party robots.txt validator before you deploy them.
Robots.txt checklist after a migration or redesign
- Confirm the live file does not contain
Disallow: /copied over from staging. - Check that CSS, JavaScript and image folders are crawlable.
- Make sure the Sitemap line uses the new domain and HTTPS.
- Verify that the file returns a 200 status. If robots.txt returns a 5xx server error for a long time, Google may stop crawling the site.
Keep a copy of every version of the file in version control. When traffic drops suddenly, a changed robots.txt is one of the first things to check.
Robots.txt vs noindex vs password protection
These three are often confused, but each does a different job:
- robots.txt Disallow: stops crawling. The URL can still appear in search if other pages link to it. Use it to save crawl effort on low-value URLs.
- noindex meta tag or header: keeps a page out of search results, but only if crawlers are allowed to fetch the page and see the tag. Don't block a noindexed page in robots.txt.
- Password protection or login: the only reliable way to keep private content away from both search engines and the public. Use it for staging sites and confidential documents.
Remember that robots.txt itself is public. Anyone can read it, so never list secret folders there expecting them to stay hidden.
