Robots.txt for AI: How to Let AI Crawlers In (Without Wrecking Your SEO)

Your robots.txt now controls access to two completely separate systems: traditional search engines and AI answer engines. Here is how to configure both without losing rankings.

By Abd Shanti Updated September 16, 2026 8 min read
XLinkedIn

The short version: Your robots.txt now controls access to two completely separate systems: traditional search engines and AI answer engines. They need separate configuration. Ignoring one means missing out on an entire traffic channel.

What robots.txt actually does

Your robots.txt file sits at the root of your website (think yoursite.com/robots.txt) and acts as a set of instructions for automated crawlers. When a bot visits your site, it checks this file first and follows the rules you have written. Simple concept, massive practical consequences.

The format was created in 1994 via the Robots Exclusion Protocol and has barely changed since. You list a User-agent name, then tell it what it can and cannot access using Allow: and Disallow: directives. If a path is not covered by any rule, crawlers can access it by default.

For years, the only crawlers that really mattered were Googlebot, Bingbot, and a handful of others. Most site owners treated robots.txt as something you set once and forgot about. That was fine in 2018. It is not fine in 2026.

Today, a new category of crawler has arrived: AI training bots and AI retrieval bots. These crawlers power products like ChatGPT, Claude, and Perplexity. Whether your content shows up in their answers depends partly on whether you have given them permission to crawl your site. And a lot of site owners have accidentally locked them out without knowing it.

Key Takeaway

Your robots.txt now controls access to two completely separate systems: traditional search engines and AI answer engines. They need separate configuration. Ignoring one means missing out on an entire traffic channel.

This article walks through every major AI crawler, the mistakes that accidentally block them, how to write a proper robots.txt that gives you control, and how to check that your settings are actually working. If you want to skip ahead and just see your current status, the AI Crawler Checker will show you which bots your current file is blocking.

How AI crawlers read robots.txt

Good news: AI crawlers follow the exact same Robots Exclusion Protocol as Google and Bing. There is no new standard to learn. They visit /robots.txt, read the file, and obey User-agent, Allow, and Disallow directives just like every other well-behaved crawler.

The key difference is the user-agent string. Googlebot identifies itself as Googlebot. OpenAI's training crawler identifies itself as GPTBot. Anthropic's crawler says ClaudeBot. If your robots.txt has a rule for User-agent: * that disallows something, that rule applies to all of them unless you have a more specific rule for that particular user-agent.

Two types of AI crawlers

It helps to understand that AI crawlers fall into two categories:

This distinction matters because blocking a training crawler means your content is not in future AI responses. Blocking a retrieval crawler means your content does not show up in current AI search answers, even if your site is excellent. Different goals, different consequences for blocking.

Key Takeaway

AI crawlers follow the same robots.txt rules as traditional search crawlers. The only thing that changes is the user-agent name. A wildcard Disallow affects them all unless you override it with bot-specific rules.

Complete list of AI bot user-agents

Here are the AI agents that matter as of September 2026. The most important column is the purpose, because AI companies now run separate agents for training, for search, and for fetching a page when a user asks. Blocking the wrong one is the most common reason a site that wants AI traffic gets none. Each name links to a full reference with rules and a live access check.

AgentCompanyPurposeFollows robots.txt
GPTBotOpenAITraining data onlyYes
OAI-SearchBotOpenAIChatGPT search indexYes
ChatGPT-UserOpenAIOpens a page when a user or GPT asksNot guaranteed
ClaudeBotAnthropicTraining dataYes
Claude-SearchBotAnthropicClaude search qualityYes
Claude-UserAnthropicOpens a page when a Claude user asksYes, per Anthropic
PerplexityBotPerplexityPerplexity search indexYes
Perplexity-UserPerplexityOpens a page when a user asksGenerally no
Google-ExtendedGoogleToken: Gemini training and grounding, not SearchToken only
Applebot-ExtendedAppleToken: Apple model trainingToken only
meta-externalagentMetaAI training and content indexingYes
AmazonbotAmazonAmazon services such as Alexa answersYes
CCBotCommon CrawlOpen dataset used by many modelsYes
BytespiderByteDanceByteDance AI trainingReported inconsistent
DuckAssistBotDuckDuckGoSources for AI-assisted answersYes

Two corrections to older lists you may find elsewhere, including an earlier version of this guide. First, GPTBot is not used for ChatGPT search. Search is OAI-SearchBot, and blocking GPTBot does not remove you from ChatGPT answers. Second, Google-Extended and Applebot-Extended are not crawlers. They are tokens that tell Google and Apple how content their normal crawlers collected may be used, so they never appear in your logs, and blocking them does not affect Google Search, AI Overviews, Siri or Spotlight.

Older tokens such as anthropic-ai and FacebookBot still appear in many robots.txt templates. They do no harm, but they are not the agents Anthropic and Meta document today. The full set of 20 agents is in the AI crawler directory, and our free AI Crawler Checker tests your live file against them.

Common mistakes that accidentally block AI bots

Most sites that are blocking AI crawlers are not doing it intentionally. They ended up that way through a plugin update, a security configuration, or a well-meaning developer who set something overly broad. Here are the most common culprits.

The wildcard Disallow everything rule

This is the most common accident. Someone pastes this into their robots.txt (often while setting up a staging environment) and never removes it:

User-agent: *
Disallow: /

This single rule blocks every crawler on the planet. Googlebot. Bingbot. GPTBot. All of them. If your site is live and has this in robots.txt, you have a serious problem that goes well beyond AI crawlers.

WordPress security plugins

Some WordPress security plugins (Wordfence, iThemes Security, and others) add crawler restrictions to your robots.txt automatically. They often block unfamiliar user-agents as a protective measure. The problem is that AI crawlers are relatively new, so many of these plugins flag them as suspicious. Check your robots.txt after installing or updating any security plugin.

Cloudflare bot fight mode and security rules

Cloudflare's Bot Fight Mode and Super Bot Fight Mode are designed to block automated traffic. Unless you have specifically whitelisted AI crawlers, they may be getting blocked at the network level, before they even reach your robots.txt. This is a separate layer of access control that robots.txt cannot override.

If you use Cloudflare, check your firewall rules and Bot Fight Mode settings. You may need to create exceptions for AI crawler IP ranges, which each company publishes (OpenAI publishes theirs at https://openai.com/gptbot-ranges.txt).

Overly broad third-party disallow rules

Some SEO tools or site migration services generate robots.txt files that block everything except Googlebot and Bingbot. This made sense before AI crawlers existed. Now it means you have explicitly told every AI bot to leave. Common pattern:

User-agent: *
Disallow: /

User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

That setup gives Google and Bing access while locking out every other crawler, including all AI bots. If you inherited this pattern, fixing it is straightforward, covered in the next section.

CDN and caching layer rewrites

If your CDN or reverse proxy serves a cached robots.txt, edits to your actual file may not propagate immediately. After making changes, verify by visiting your live /robots.txt URL directly in a browser, not just looking at the source file on your server.

Key Takeaway

Most accidental AI bot blocks come from wildcard Disallow rules, security plugins, or Cloudflare settings. Check your live robots.txt URL in a browser right now and scan it with the AI Crawler Checker to see your actual status.

How to allow specific AI bots

Robots.txt works in groups. A group starts with one or more User-agent lines followed by its rules. Each crawler looks for the group whose user agent matches its own name most specifically and follows only that group. Two consequences trip up almost everyone:

Allow selected crawlers while blocking everyone else

User-agent: *
Disallow: /

User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /admin/
Disallow: /checkout/
Allow: /

Listing several User-agent lines above one set of rules gives all of those crawlers the same group, which keeps the file short and avoids forgetting a Disallow line in one of them.

Stay in AI search, opt out of AI training

User-agent: *
Disallow: /admin/
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /

This keeps OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot and Bingbot on the permissive wildcard group, so you can still be found and cited, while training crawlers and training tokens are told no.

How to allow all AI bots at once

If your goal is maximum AI visibility and you want every legitimate crawler to have full access, the simplest approach is a universal allow. This is the configuration most content sites and publishers should start with.

User-agent: *
Allow: /

That is the entire file. Two lines. The wildcard * applies to every crawler, and Allow: / means the entire site is accessible. No explicit Disallow means everything is open by default anyway, so this is technically equivalent to an empty robots.txt, but it is explicit and clear about your intent.

You can still exclude specific paths you do not want crawled. Admin areas, user dashboards, and checkout pages are common things to exclude even when you want AI crawlers to access your content:

User-agent: *
Allow: /
Disallow: /admin/
Disallow: /dashboard/
Disallow: /checkout/
Disallow: /account/

This gives all crawlers access to your public content while keeping sensitive areas off-limits. For most content-focused sites, this is the right balance.

The blocking debate: reasons to allow vs. block

Plenty of site owners have real reasons to think carefully about whether to allow AI crawlers. This is not a simple "allow everything" situation for everyone. Here are the genuine arguments on both sides.

Reasons to allow AI crawlers

Reasons to block AI crawlers

There is no universally correct answer. The right choice depends on your business model, your content type, and your views on how AI platforms relate to your industry. What matters is making an intentional decision rather than an accidental one.

For a broader look at how AI search works and why this all matters for your site, the What is AI SEO article covers the full picture.

Full example robots.txt for maximum AI access

Here is a production-ready robots.txt that lets every major AI and search crawler read your public content while keeping sensitive paths out. It deliberately uses a single wildcard group for the paths, because a separate Allow: / group per bot would silently give those bots access to everything the wildcard disallows.

# robots.txt: maximum AI crawler access
# Last updated: 2026-09-16

# One group for every crawler, so the Disallow lines apply to all of them.
User-agent: *
Disallow: /admin/
Disallow: /account/
Disallow: /checkout/
Disallow: /cart/
Disallow: /login/
Disallow: /wp-admin/
Disallow: /?s=
Disallow: /search
Allow: /wp-admin/admin-ajax.php
Allow: /

# Optional: keep a crawler you do not want out entirely.
# A named group replaces the wildcard for that bot, so it only needs its own rules.
User-agent: Bytespider
Disallow: /

Sitemap: https://yoursite.com/sitemap.xml

A few notes on this configuration. You do not need to list GPTBot, OAI-SearchBot, ClaudeBot or the others by name to allow them: the wildcard group already does, and adding a named Allow: / group for each would remove your Disallow lines for those bots. Robots.txt also cannot override a firewall or CDN. If Cloudflare, a security plugin or your host blocks AI bots, they are refused before they ever read this file, so check those layers separately. Keep the Sitemap line pointing at your real sitemap so crawlers can discover every page.

You can also build a file like this with the Robots.txt Generator.

How to verify your settings actually work

Writing the robots.txt is step one. Confirming that crawlers can read it, and are not blocked somewhere else, is step two.

1. Check the live file

Open https://yoursite.com/robots.txt in a browser. What you see is what crawlers see. If it is blank, returns an error, or shows an old version, fix the serving or cache issue first.

2. Use the robots.txt report in Google Search Console

Google retired its old robots.txt Tester in 2023. The replacement is the robots.txt report in Search Console settings, which shows the robots.txt files Google found for your site, when they were last crawled, and any warnings or errors. To test whether Googlebot may crawl a specific URL, use URL Inspection.

3. Test each AI agent against your file

Our free AI Crawler Checker reads your live robots.txt and shows the allow or block status for the major AI agents at once. Every page in the AI crawler directory also has a live check for that single agent, including which group and which rule decided the result.

4. Look for blocks outside robots.txt

Bot protection can refuse crawlers that robots.txt allows. The ChatGPT visibility checker fetches your page as well as your robots.txt, so it catches a firewall or challenge page that a robots.txt test alone would miss. Your server logs are the final word: if an agent is allowed but never appears, something upstream is stopping it.

Re-crawl timing

Crawlers pick up robots.txt changes on their next visit. OpenAI says its search systems adjust to robots.txt changes in about a day. For Bing, which several AI assistants rely on, you can speed up discovery of changed pages with IndexNow. Other crawlers may take longer, so give changes a few days before drawing conclusions.

Frequently Asked Questions

Blocking GPTBot stops OpenAI from collecting your pages as training data from then on. It does not remove you from ChatGPT search answers, which rely on OAI-SearchBot and on live fetches by ChatGPT-User. To stay out of ChatGPT search as well, you would need to block OAI-SearchBot. Either change is reversible by editing robots.txt.
Partly. Blocking GPTBot and ClaudeBot should stop those companies from using your content for future training. But Common Crawl (CCBot) has already archived much of the web, and many models trained on those archives before crawler blocking was common. Blocking crawlers protects future training, not past usage.
No. Google's crawlers (Googlebot, AdsBot) are completely separate from AI crawlers like GPTBot and ClaudeBot. Allowing or blocking AI crawlers has zero effect on Googlebot's access or your Google rankings. They are independent systems.
Almost certainly not. <code>Disallow: /</code> blocks ALL crawlers including Googlebot and Bingbot. This makes your site invisible to search engines and kills organic traffic entirely. If you want to block specific crawlers, list them explicitly by User-agent name.
Visit <code>yoursite.com/robots.txt</code> in a browser to see your current file. Or use our free <a href="/tools/ai-crawler-checker">AI Crawler Checker</a> at freegptseo.com/tools/ai-crawler-checker which tests your robots.txt against every major AI bot user-agent and shows you exactly which ones are blocked.
Abd Shanti

Abd Shanti

Co-founder and strategy lead at Outline Technologies, the team behind FreeGPTSEO and AI Citation Monitor. Abd works on how brands get found and cited by search engines and AI assistants.

Check How You Score Right Now

Run a free AI SEO audit on your site. See your score across schema, content, meta tags, and AI crawler access. Takes 5 seconds.

Run Free Audit
Last updated: September 16, 2026. Agent roster and robots.txt behaviour re-checked against each operator's documentation.