Get in touch

Robots.txt: Complete Guide and Best Practices

Robots.txt · · By Monika Gupta
Robots.txt: Complete Guide and Best Practices

The robots.txt file is a plain text file at the root of your domain (e.g., yourdomain.com/robots.txt) that tells search engine crawlers which parts of your site they’re allowed — or not allowed — to access.

Basic syntax: Each rule specifies a User-agent (which crawler it applies to) and Disallow or Allow directives (which paths to block or permit). For example, blocking a staging folder from all crawlers looks like:

User-agent: *
Disallow: /staging/

Common mistakes: The most damaging error is accidentally blocking your entire site with Disallow: / after a site migration or redesign — this tells every crawler to stay away from everything. Another frequent issue is blocking CSS or JavaScript files that Google needs to render the page properly, which can hurt how your content is understood.

What robots.txt does NOT do: It doesn’t remove pages from Google’s index — a blocked page can still appear in search results (usually without a description) if other sites link to it. To fully keep a page out of search results, use a noindex meta tag instead, on a page that robots.txt allows crawlers to reach.

Best practices: Keep the file simple, always include your sitemap URL at the bottom, check the robots.txt report in Search Console after deploying, and audit it after every major site migration.

A practical robots.txt example

Here is a typical file for a small business site or online store:

User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /checkout/
Disallow: /*?sessionid=

Sitemap: https://example.com/sitemap.xml

This keeps crawlers out of private and low-value areas, blocks session-ID URLs that create endless duplicates, and points every crawler to the sitemap.

Wildcards and pattern matching

Google and Bing support two special characters:

  • * matches any sequence of characters. Disallow: /*?filter= blocks every URL containing ?filter=.
  • $ marks the end of a URL. Disallow: /*.pdf$ blocks URLs that end in .pdf, but not a page like /pdf-guide/.

When an Allow and a Disallow rule both match a URL, Google follows the most specific (longest) rule. This lets you block a folder but allow one file inside it.

Controlling AI crawlers

Robots.txt is also how you decide whether AI crawlers can use your content. Common user-agents include GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot and Google-Extended, which controls whether Google may use your content for Gemini training. Google-Extended does not affect Google Search or AI Overviews; those use the normal Googlebot. Before blocking AI crawlers, decide whether being cited in AI answers matters to your business. For most brands that want visibility in ChatGPT and Perplexity, blocking them works against that goal.

How to test robots.txt safely

Google retired the old robots.txt Tester in 2023. Today you can:

  1. Open the robots.txt report in Search Console (Settings → robots.txt) to see the version Google last fetched and any parsing errors.
  2. Use URL Inspection on an important page to confirm it is not blocked.
  3. Test new rules in a third-party robots.txt validator before you deploy them.

Robots.txt checklist after a migration or redesign

  • Confirm the live file does not contain Disallow: / copied over from staging.
  • Check that CSS, JavaScript and image folders are crawlable.
  • Make sure the Sitemap line uses the new domain and HTTPS.
  • Verify that the file returns a 200 status. If robots.txt returns a 5xx server error for a long time, Google may stop crawling the site.

Keep a copy of every version of the file in version control. When traffic drops suddenly, a changed robots.txt is one of the first things to check.

Robots.txt vs noindex vs password protection

These three are often confused, but each does a different job:

  • robots.txt Disallow: stops crawling. The URL can still appear in search if other pages link to it. Use it to save crawl effort on low-value URLs.
  • noindex meta tag or header: keeps a page out of search results, but only if crawlers are allowed to fetch the page and see the tag. Don't block a noindexed page in robots.txt.
  • Password protection or login: the only reliable way to keep private content away from both search engines and the public. Use it for staging sites and confidential documents.

Remember that robots.txt itself is public. Anyone can read it, so never list secret folders there expecting them to stay hidden.

Related Guides

Frequently asked

Not reliably — use a noindex tag for that instead.

At the root of the domain, e.g., yourdomain.com/robots.txt — subfolder locations won’t be recognized.

Yes — a misplaced Disallow: / blocks all crawling, so always test changes before publishing.

Monika Gupta

Digital Marketring Executive

ClicZeo Editorial Team is a team of digital marketing professionals specializing in SEO, AEO (Answer Engine Optimization), GEO, Google Ads, Meta Ads, content marketing, and local business growth. We create data-driven content to help businesses improve their online visibility, generate qualified leads, and stay ahead of the latest digital marketing trends.

See How AI Search Describes Your Brand Today

We run your domain through the same visibility checks we use on client accounts — AI answer coverage, technical SEO and content gaps — and send you the findings. No obligation.