What's in This Guide
- What robots.txt actually does
- How AI crawlers read robots.txt
- Complete list of AI bot user-agents
- Common mistakes that accidentally block AI bots
- How to allow specific AI bots
- How to allow all AI bots at once
- The blocking debate: reasons to allow vs. block
- Full example robots.txt for maximum AI access
- How to verify your settings actually work
- Frequently asked questions
The short version: Your robots.txt now controls access to two completely separate systems: traditional search engines and AI answer engines. They need separate configuration. Ignoring one means missing out on an entire traffic channel.
What robots.txt actually does
Your robots.txt file sits at the root of your website (think yoursite.com/robots.txt) and acts as a set of instructions for automated crawlers. When a bot visits your site, it checks this file first and follows the rules you have written. Simple concept, massive practical consequences.
The format was created in 1994 via the Robots Exclusion Protocol and has barely changed since. You list a User-agent name, then tell it what it can and cannot access using Allow: and Disallow: directives. If a path is not covered by any rule, crawlers can access it by default.
For years, the only crawlers that really mattered were Googlebot, Bingbot, and a handful of others. Most site owners treated robots.txt as something you set once and forgot about. That was fine in 2018. It is not fine in 2026.
Today, a new category of crawler has arrived: AI training bots and AI retrieval bots. These crawlers power products like ChatGPT, Claude, and Perplexity. Whether your content shows up in their answers depends partly on whether you have given them permission to crawl your site. And a lot of site owners have accidentally locked them out without knowing it.
Key TakeawayYour robots.txt now controls access to two completely separate systems: traditional search engines and AI answer engines. They need separate configuration. Ignoring one means missing out on an entire traffic channel.
This article walks through every major AI crawler, the mistakes that accidentally block them, how to write a proper robots.txt that gives you control, and how to check that your settings are actually working. If you want to skip ahead and just see your current status, the AI Crawler Checker will show you which bots your current file is blocking.
How AI crawlers read robots.txt
Good news: AI crawlers follow the exact same Robots Exclusion Protocol as Google and Bing. There is no new standard to learn. They visit /robots.txt, read the file, and obey User-agent, Allow, and Disallow directives just like every other well-behaved crawler.
The key difference is the user-agent string. Googlebot identifies itself as Googlebot. OpenAI's training crawler identifies itself as GPTBot. Anthropic's crawler says ClaudeBot. If your robots.txt has a rule for User-agent: * that disallows something, that rule applies to all of them unless you have a more specific rule for that particular user-agent.
Two types of AI crawlers
It helps to understand that AI crawlers fall into two categories:
- Training crawlers collect content to train AI models. They visit your site, store the content, and it may end up in a future model's training data. GPTBot and ClaudeBot are in this group.
- Retrieval crawlers fetch content in real time when a user asks a question. They want to include your content in an answer right now. OAI-SearchBot (for ChatGPT search) and PerplexityBot are in this group.
This distinction matters because blocking a training crawler means your content is not in future AI responses. Blocking a retrieval crawler means your content does not show up in current AI search answers, even if your site is excellent. Different goals, different consequences for blocking.
Key TakeawayAI crawlers follow the same robots.txt rules as traditional search crawlers. The only thing that changes is the user-agent name. A wildcard Disallow affects them all unless you override it with bot-specific rules.
Complete list of AI bot user-agents
Here are the AI agents that matter as of September 2026. The most important column is the purpose, because AI companies now run separate agents for training, for search, and for fetching a page when a user asks. Blocking the wrong one is the most common reason a site that wants AI traffic gets none. Each name links to a full reference with rules and a live access check.
| Agent | Company | Purpose | Follows robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Training data only | Yes |
| OAI-SearchBot | OpenAI | ChatGPT search index | Yes |
| ChatGPT-User | OpenAI | Opens a page when a user or GPT asks | Not guaranteed |
| ClaudeBot | Anthropic | Training data | Yes |
| Claude-SearchBot | Anthropic | Claude search quality | Yes |
| Claude-User | Anthropic | Opens a page when a Claude user asks | Yes, per Anthropic |
| PerplexityBot | Perplexity | Perplexity search index | Yes |
| Perplexity-User | Perplexity | Opens a page when a user asks | Generally no |
| Google-Extended | Token: Gemini training and grounding, not Search | Token only | |
| Applebot-Extended | Apple | Token: Apple model training | Token only |
| meta-externalagent | Meta | AI training and content indexing | Yes |
| Amazonbot | Amazon | Amazon services such as Alexa answers | Yes |
| CCBot | Common Crawl | Open dataset used by many models | Yes |
| Bytespider | ByteDance | ByteDance AI training | Reported inconsistent |
| DuckAssistBot | DuckDuckGo | Sources for AI-assisted answers | Yes |
Two corrections to older lists you may find elsewhere, including an earlier version of this guide. First, GPTBot is not used for ChatGPT search. Search is OAI-SearchBot, and blocking GPTBot does not remove you from ChatGPT answers. Second, Google-Extended and Applebot-Extended are not crawlers. They are tokens that tell Google and Apple how content their normal crawlers collected may be used, so they never appear in your logs, and blocking them does not affect Google Search, AI Overviews, Siri or Spotlight.
Older tokens such as anthropic-ai and FacebookBot still appear in many robots.txt templates. They do no harm, but they are not the agents Anthropic and Meta document today. The full set of 20 agents is in the AI crawler directory, and our free AI Crawler Checker tests your live file against them.
Common mistakes that accidentally block AI bots
Most sites that are blocking AI crawlers are not doing it intentionally. They ended up that way through a plugin update, a security configuration, or a well-meaning developer who set something overly broad. Here are the most common culprits.
The wildcard Disallow everything rule
This is the most common accident. Someone pastes this into their robots.txt (often while setting up a staging environment) and never removes it:
User-agent: *Disallow: /
This single rule blocks every crawler on the planet. Googlebot. Bingbot. GPTBot. All of them. If your site is live and has this in robots.txt, you have a serious problem that goes well beyond AI crawlers.
WordPress security plugins
Some WordPress security plugins (Wordfence, iThemes Security, and others) add crawler restrictions to your robots.txt automatically. They often block unfamiliar user-agents as a protective measure. The problem is that AI crawlers are relatively new, so many of these plugins flag them as suspicious. Check your robots.txt after installing or updating any security plugin.
Cloudflare bot fight mode and security rules
Cloudflare's Bot Fight Mode and Super Bot Fight Mode are designed to block automated traffic. Unless you have specifically whitelisted AI crawlers, they may be getting blocked at the network level, before they even reach your robots.txt. This is a separate layer of access control that robots.txt cannot override.
If you use Cloudflare, check your firewall rules and Bot Fight Mode settings. You may need to create exceptions for AI crawler IP ranges, which each company publishes (OpenAI publishes theirs at https://openai.com/gptbot-ranges.txt).
Overly broad third-party disallow rules
Some SEO tools or site migration services generate robots.txt files that block everything except Googlebot and Bingbot. This made sense before AI crawlers existed. Now it means you have explicitly told every AI bot to leave. Common pattern:
User-agent: *Disallow: /
User-agent: GooglebotAllow: /
User-agent: BingbotAllow: /
That setup gives Google and Bing access while locking out every other crawler, including all AI bots. If you inherited this pattern, fixing it is straightforward, covered in the next section.
CDN and caching layer rewrites
If your CDN or reverse proxy serves a cached robots.txt, edits to your actual file may not propagate immediately. After making changes, verify by visiting your live /robots.txt URL directly in a browser, not just looking at the source file on your server.
Key TakeawayMost accidental AI bot blocks come from wildcard Disallow rules, security plugins, or Cloudflare settings. Check your live robots.txt URL in a browser right now and scan it with the AI Crawler Checker to see your actual status.
How to allow specific AI bots
Robots.txt works in groups. A group starts with one or more User-agent lines followed by its rules. Each crawler looks for the group whose user agent matches its own name most specifically and follows only that group. Two consequences trip up almost everyone:
- Order does not matter. A crawler uses the most specific matching group wherever it sits in the file. Moving a group above or below
User-agent: *changes nothing. - A named group replaces the wildcard. Once GPTBot has its own group, it ignores every rule under
User-agent: *. If your wildcard group disallows/admin/, GPTBot is no longer told to stay out of/admin/unless you repeat that line in its group.
Allow selected crawlers while blocking everyone else
User-agent: *
Disallow: /
User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /admin/
Disallow: /checkout/
Allow: /
Listing several User-agent lines above one set of rules gives all of those crawlers the same group, which keeps the file short and avoids forgetting a Disallow line in one of them.
Stay in AI search, opt out of AI training
User-agent: *
Disallow: /admin/
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /
This keeps OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot and Bingbot on the permissive wildcard group, so you can still be found and cited, while training crawlers and training tokens are told no.
How to allow all AI bots at once
If your goal is maximum AI visibility and you want every legitimate crawler to have full access, the simplest approach is a universal allow. This is the configuration most content sites and publishers should start with.
User-agent: *Allow: /
That is the entire file. Two lines. The wildcard * applies to every crawler, and Allow: / means the entire site is accessible. No explicit Disallow means everything is open by default anyway, so this is technically equivalent to an empty robots.txt, but it is explicit and clear about your intent.
You can still exclude specific paths you do not want crawled. Admin areas, user dashboards, and checkout pages are common things to exclude even when you want AI crawlers to access your content:
User-agent: *Allow: /Disallow: /admin/Disallow: /dashboard/Disallow: /checkout/Disallow: /account/
This gives all crawlers access to your public content while keeping sensitive areas off-limits. For most content-focused sites, this is the right balance.
The blocking debate: reasons to allow vs. block
Plenty of site owners have real reasons to think carefully about whether to allow AI crawlers. This is not a simple "allow everything" situation for everyone. Here are the genuine arguments on both sides.
Reasons to allow AI crawlers
- AI citation traffic is growing. When ChatGPT or Perplexity cites your site in an answer, users click through. This is a real and growing source of referral traffic. Blocking AI bots means opting out of that channel entirely.
- Brand awareness in AI answers. Even when users do not click, seeing your brand name cited in an AI response builds recognition and trust. Publishers who block AI crawlers become invisible in these conversations.
- AI SEO is still early. Sites that establish a strong presence in AI training data and retrieval systems now are likely to benefit as these platforms grow. Early access creates a compounding advantage over time.
- Retrieval-only bots do not use your content for training. If your concern is about training data, you can block training crawlers (like CCBot) while still allowing retrieval-only crawlers (like OAI-SearchBot and PerplexityBot) that fetch your content to cite in live answers.
Reasons to block AI crawlers
- Copyright and content ownership. Many publishers believe that AI companies are commercially benefiting from their content without compensation. Blocking crawlers is currently the primary mechanism available to opt out of this.
- Scraping without attribution. Some AI products summarize content without linking back to the source, eliminating the traffic benefit and leaving only the cost of being crawled.
- Server load. AI training crawlers can be aggressive and consume significant bandwidth. For small sites or those with limited hosting resources, blocking aggressive crawlers is a practical operational decision.
- Legal uncertainty. The legal status of using web content for AI training is still being decided in courts globally. Some publishers prefer to opt out now rather than be included in training datasets while the legal landscape is unclear.
There is no universally correct answer. The right choice depends on your business model, your content type, and your views on how AI platforms relate to your industry. What matters is making an intentional decision rather than an accidental one.
For a broader look at how AI search works and why this all matters for your site, the What is AI SEO article covers the full picture.
Full example robots.txt for maximum AI access
Here is a production-ready robots.txt that lets every major AI and search crawler read your public content while keeping sensitive paths out. It deliberately uses a single wildcard group for the paths, because a separate Allow: / group per bot would silently give those bots access to everything the wildcard disallows.
# robots.txt: maximum AI crawler access
# Last updated: 2026-09-16
# One group for every crawler, so the Disallow lines apply to all of them.
User-agent: *
Disallow: /admin/
Disallow: /account/
Disallow: /checkout/
Disallow: /cart/
Disallow: /login/
Disallow: /wp-admin/
Disallow: /?s=
Disallow: /search
Allow: /wp-admin/admin-ajax.php
Allow: /
# Optional: keep a crawler you do not want out entirely.
# A named group replaces the wildcard for that bot, so it only needs its own rules.
User-agent: Bytespider
Disallow: /
Sitemap: https://yoursite.com/sitemap.xml
A few notes on this configuration. You do not need to list GPTBot, OAI-SearchBot, ClaudeBot or the others by name to allow them: the wildcard group already does, and adding a named Allow: / group for each would remove your Disallow lines for those bots. Robots.txt also cannot override a firewall or CDN. If Cloudflare, a security plugin or your host blocks AI bots, they are refused before they ever read this file, so check those layers separately. Keep the Sitemap line pointing at your real sitemap so crawlers can discover every page.
You can also build a file like this with the Robots.txt Generator.
How to verify your settings actually work
Writing the robots.txt is step one. Confirming that crawlers can read it, and are not blocked somewhere else, is step two.
1. Check the live file
Open https://yoursite.com/robots.txt in a browser. What you see is what crawlers see. If it is blank, returns an error, or shows an old version, fix the serving or cache issue first.
2. Use the robots.txt report in Google Search Console
Google retired its old robots.txt Tester in 2023. The replacement is the robots.txt report in Search Console settings, which shows the robots.txt files Google found for your site, when they were last crawled, and any warnings or errors. To test whether Googlebot may crawl a specific URL, use URL Inspection.
3. Test each AI agent against your file
Our free AI Crawler Checker reads your live robots.txt and shows the allow or block status for the major AI agents at once. Every page in the AI crawler directory also has a live check for that single agent, including which group and which rule decided the result.
4. Look for blocks outside robots.txt
Bot protection can refuse crawlers that robots.txt allows. The ChatGPT visibility checker fetches your page as well as your robots.txt, so it catches a firewall or challenge page that a robots.txt test alone would miss. Your server logs are the final word: if an agent is allowed but never appears, something upstream is stopping it.
Re-crawl timing
Crawlers pick up robots.txt changes on their next visit. OpenAI says its search systems adjust to robots.txt changes in about a day. For Bing, which several AI assistants rely on, you can speed up discovery of changed pages with IndexNow. Other crawlers may take longer, so give changes a few days before drawing conclusions.
Frequently Asked Questions
Check How You Score Right Now
Run a free AI SEO audit on your site. See your score across schema, content, meta tags, and AI crawler access. Takes 5 seconds.
Run Free Audit