GEO · Crawler Access

OAI-SearchBot vs GPTBot: The robots.txt Decision Most Brands Get Wrong

Two identical metal doors side by side on a dark wall, one standing open with warm light inside and the other closed with a padlock, standing in for two crawlers with opposite robots.txt settings.

OAI-SearchBot decides whether your pages can appear in ChatGPT search answers. GPTBot decides whether your content can be used to train OpenAI's foundation models. They are separate robots.txt rules, so a site can allow one and block the other. One blanket rule that blocks both is the mistake that quietly removes a brand from ChatGPT search.

OpenAI, Anthropic, Perplexity and Google each split their crawlers by purpose in their own way. The decision is small to make and easy to get backwards, because the names look alike and many robots.txt templates lump them together. Choosing per crawler takes a few lines and a check against the live file.

Key Takeaways

  • OAI-SearchBot surfaces sites in ChatGPT search; GPTBot collects content for model training; ChatGPT-User fetches pages when a person asks
  • OpenAI states that each setting is independent of the others
  • Blocking OAI-SearchBot removes a site from ChatGPT search answers; blocking GPTBot addresses training
  • ChatGPT-User is user-initiated, so robots.txt rules may not apply to it
  • Google-Extended is a robots.txt token that does not affect Search inclusion or ranking
  • Allowing a crawler makes a page eligible to be cited, and the setting does not decide whether ChatGPT cites it

What is the difference between OAI-SearchBot and GPTBot?

OpenAI documents three separate crawlers, each with its own purpose and its own robots.txt setting. OAI-SearchBot is used to surface websites in search results in ChatGPT's search features, and OpenAI states that a site that opts out of it will not appear in ChatGPT search answers. GPTBot is used to crawl content that helps make OpenAI's generative AI foundation models more useful and safe, and disallowing it indicates that a site's content should not be used in training those models. OpenAI adds that each setting is independent of the others, which is the sentence the rest of this article depends on. Source: OpenAI's crawler documentation, accessed September 2026.

ChatGPT-User is the third agent, and it works differently from the other two. OpenAI describes it as used for certain user actions in ChatGPT and Custom GPTs, which means it visits a page because a person asked for that page to be read. Because the visit is user-initiated, OpenAI notes that robots.txt rules may not apply to it. A robots.txt file therefore controls the two automated crawlers well and controls ChatGPT-User only partly. OpenAI publishes IP ranges for each of the three agents, so a visit in your server logs can be verified against the list instead of trusted on the strength of a user-agent string. Source: the same OpenAI documentation, accessed September 2026.

Which robots.txt setup fits which goal?

The right setup depends on two separate questions. The first is whether the brand wants to be eligible for ChatGPT search answers, which is a visibility question and belongs to marketing. The second is whether the brand is comfortable with its content being used to train models, which is a rights question and belongs to whoever owns the content and, where contracts or licensing are involved, to legal counsel. The table below sets the four possible combinations side by side. Each row describes what the two rules do according to OpenAI's documentation, accessed September 2026, and says nothing about how ChatGPT then ranks or cites a page.

GoalOAI-SearchBotGPTBotWhat happens
Eligible for ChatGPT search, training use acceptedAllowAllowPages can surface in ChatGPT search and can be used for training
Eligible for ChatGPT search, training use declinedAllowDisallowPages can surface in search; content is signalled as off limits for training
Out of bothDisallowDisallowSite is absent from ChatGPT search answers and signalled as off limits for training
Search blocked, training allowedDisallowAllowSite is absent from ChatGPT search answers while training use stays open; rarely intentional

Rows one and two both keep the brand eligible for ChatGPT search. Eligibility is what a crawler rule delivers: allowing OAI-SearchBot lets the crawler read the page, and the rule does not decide whether ChatGPT cites that page in an answer. Nobody credible sells guaranteed citations, and a robots.txt line cannot create one. For a brand that wants training use declined and search visibility kept, the setup looks like this:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Why does a blanket AI bot rule take a brand out of ChatGPT search?

Google's robots.txt specification explains the matching logic that crawlers follow: only one group of rules is valid for a particular crawler, and the crawler picks the group with the most specific user agent that matches it. Rules from a specific group and from the global group are not combined. In practice, a crawler that is named in the file follows its own group and ignores the rest, while a crawler that is not named falls back to the wildcard group. Google's page describes Google's own crawlers, and OpenAI's documentation does not restate the matching rule, so treat its application to OpenAI's bots as the standard convention and confirm it with a test. Source: Google Search Central, accessed September 2026.

Three patterns cause most of the accidental blocks. The first is a copied template that lists every AI crawler in one group with a Disallow line, which puts OAI-SearchBot next to GPTBot and treats them as the same thing. The second is a wildcard Disallow left over from a staging or launch period, which blocks any crawler that has no group of its own. The third sits outside the file entirely: firewall and bot-management settings can refuse a request that robots.txt allows, because robots.txt states a preference and does not enforce anything. Google's specification is explicit that robots.txt is not an access control mechanism, so private material belongs behind authentication.

We keep a running set of plain answers to the questions clients ask about AI search visibility, each checked against official documentation.

Read the answers

Do Google, Anthropic and Perplexity split search and training the same way?

Each provider draws the line differently, and the differences matter when a single robots.txt file has to cover all of them. The table lists what each provider's own documentation says, accessed September 2026. A column can say what a provider documents about its crawlers, and it cannot say how the provider's product then treats a given site, so the entries are kept narrow on purpose. Where a provider documents no separate training crawler, the table says so instead of guessing.

ProviderSearch or answer crawlerTraining crawlerUser-initiated fetch
OpenAIOAI-SearchBotGPTBotChatGPT-User; robots.txt rules may not apply
AnthropicClaude-SearchBotClaudeBotClaude-User; honors robots.txt exclusions
PerplexityPerplexityBotNone: documented as not used for foundation modelsPerplexity-User; generally ignores robots.txt
GoogleGooglebot for SearchGoogle-Extended token for GeminiNot covered here

Anthropic's help article lists three bots. ClaudeBot collects web content to improve its models, Claude-User lets Claude access websites during conversations and web searches, and Claude-SearchBot indexes content to improve search result quality, with all three respecting robots.txt directives. Perplexity's documentation says PerplexityBot surfaces and links websites in Perplexity search and is not used for AI foundation models, while Perplexity-User visits a page when a person asks and generally ignores robots.txt because the user requested the fetch. Google's crawler documentation says Google-Extended governs whether crawled content may train future Gemini models and ground Gemini apps, has no separate user-agent string, and does not affect inclusion in Google Search or act as a ranking signal.

How do I check that my robots.txt is doing what I intend?

Start with the live file, not the copy in the repository or the CMS. Fetch it from the public URL, confirm that each crawler you care about has a group with its own name, and read the rules under that name as the crawler would. OpenAI states that it can take around 24 hours from a robots.txt update for its systems to adjust, so a change made today is not a fair test until tomorrow. After that window, check server logs for the OAI-SearchBot user agent, match the source addresses against the IP list OpenAI publishes, and look for 403 or 429 responses that would point to a firewall rule instead of the file. Source: OpenAI's crawler documentation, accessed September 2026.

Our own file is a working example. Cipion's robots.txt names OAI-SearchBot, GPTBot and ChatGPT-User in three separate groups, all set to Allow, alongside the equivalent groups for Anthropic, Perplexity and Google, and it was last reviewed on 18 September 2026. Before that review the file named GPTBot but did not name OAI-SearchBot, so it said nothing explicit about the crawler that feeds ChatGPT search. Nothing about that change is a ranking result, and we make no claim that it has moved a citation. It is a hygiene fix, and the kind of check we describe in our guide to programming an agent for GEO maintenance.

Common questions

Should I block GPTBot?

That is a decision about training rights, and it does not depend on search visibility. According to OpenAI, disallowing GPTBot indicates that a site's content should not be used to train its generative AI foundation models, while OAI-SearchBot is a separate setting that governs appearance in ChatGPT search. A brand can disallow GPTBot and still allow OAI-SearchBot. Whether to block belongs to whoever owns the content, and to legal counsel where contracts or licensing apply.

If I allow OAI-SearchBot, will ChatGPT cite my site?

Not by itself. Allowing OAI-SearchBot makes pages eligible to be surfaced in ChatGPT search features, and blocking it removes the site from ChatGPT search answers. Eligibility is a precondition and a placement is something different. Which pages get cited depends on the question asked and on the content available, and no crawler setting can promise a citation.

How long does a robots.txt change take to apply to OpenAI's crawlers?

OpenAI states that it can take around 24 hours from a site's robots.txt update for its systems to adjust. After changing the file, wait at least a day before deciding whether the change worked, and verify the live file at the public URL rather than the version in a CMS or repository.

Does blocking ChatGPT-User in robots.txt stop ChatGPT from reading my page?

Not reliably. OpenAI describes ChatGPT-User as used for certain user actions in ChatGPT and Custom GPTs, and notes that because those actions are initiated by a user, robots.txt rules may not apply. Content that must stay private should sit behind authentication, since robots.txt is not an access control mechanism.

Does Google-Extended affect my Google rankings?

No. Google's documentation states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal. It controls whether crawled content may be used to train future Gemini models and for grounding in Gemini apps, and it has no separate HTTP user-agent string of its own.

Is robots.txt enough to protect private content?

No. Google's robots.txt documentation says the file is not an access control mechanism, and a disallowed URL can still be indexed without a snippet if it is discovered elsewhere. Private pages need HTTP authentication or an equivalent control, with robots.txt used to express crawling preferences for public content.

Not sure what your robots.txt says to AI crawlers?

We read the live file, your server responses and your firewall rules together and tell you which crawlers can reach which pages. No guaranteed citations, because nobody can sell those.

See our GEO service