GuidesBy 4 min read

Is your website blocking ChatGPT without you knowing? A guide to AI crawlers

GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot: what each AI crawler does, how to set up robots.txt and what Cloudflare's new rules change for your site.

Cover: Is your website blocking ChatGPT without you knowing? A guide to AI crawlers

In short

ChatGPT, Claude and Perplexity can only cite you if their search crawlers can read your site. You can block model training and stay citable, but you need to tell the right crawlers apart in robots.txt and check your CDN and firewall too, starting with Cloudflare's new settings of 15 September 2026.

  • Every AI provider uses different crawlers for search, for training and for user requests.
  • Blocking GPTBot or ClaudeBot stops training; blocking OAI-SearchBot or Claude-SearchBot removes you from answers.
  • Since 15 September 2026 Cloudflare blocks Training and Agent by default on ad pages of new domains.
  • The check takes ten minutes: robots.txt, CDN settings, content in the HTML and server logs.

Written by the Geosnap AI agent, reviewed and approved by Rinald Sefa.

Translated from the Italian original. Read the original

If ChatGPT, Claude or Perplexity never cite your website, the cause may be more mundane than you think: a rule in your robots.txt file, a CDN setting or a firewall that keeps their crawlers out. Nobody notices, because people still see the site as usual. This guide covers which AI crawlers exist, what each one does, how to decide which ones to let in, and what changed with Cloudflare's new settings.

Why robots.txt matters for getting cited by AI

AI assistants can only cite pages they can read. Before answering a question, ChatGPT, Claude and Perplexity search the web with their own crawlers: if your site blocks them, someone else shows up in the answer.

robots.txt is the file at the root of your site that tells crawlers what they may visit. The AI crawlers of the major providers say they respect it, so a single Disallow line is enough to make you disappear from their searches. The problem is that many of these lines were added in recent years to block model training, without distinguishing between crawlers that collect training data and crawlers looking for sources to cite.

Which AI crawlers exist and what they do

Each provider runs several crawlers with different purposes. These are the officially documented ones:

ProviderUser agentPurposeRespects robots.txt
OpenAIOAI-SearchBotFinding sites to show in ChatGPT searchYes
OpenAIGPTBotCollecting content to train modelsYes
OpenAIChatGPT-UserVisiting a page when a user asksNot always
AnthropicClaude-SearchBotIndexing content for Claude's searchesYes
AnthropicClaudeBotCollecting content for trainingYes
AnthropicClaude-UserReading a page at a user's requestYes
PerplexityPerplexityBotSurfacing and linking sites in Perplexity resultsYes
PerplexityPerplexity-UserVisiting a page at a user's requestGenerally no
GoogleGoogle-ExtendedUse of content to train Gemini and for groundingYes

OpenAI explains that sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, at most as navigational links. Perplexity notes that Perplexity-User generally ignores robots.txt, because a person requested the visit. Google, for its part, states that Google-Extended does not affect inclusion in Google Search or ranking.

Search and training can be separated

Yes: you can keep your content out of model training and still be citable. OpenAI says so explicitly: each setting is independent of the others, so you can allow OAI-SearchBot and block GPTBot. The same goes for Anthropic's crawlers: according to Claude's documentation, blocking ClaudeBot excludes future content from training, while blocking Claude-SearchBot may reduce your visibility in results.

A robots.txt that stays visible in answers without contributing to training can look like this:

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

It's a choice, not a rule: many brands prefer to leave training open too, because models learn about the brand. What matters is that it's a deliberate decision, not a side effect.

Cloudflare: what changed on 15 September 2026

If your site runs through Cloudflare, check its bot settings: CDN rules apply even when robots.txt is open. In July 2026 Cloudflare split AI crawlers into three categories: Search (indexing for search results), Agent (automated actions on behalf of users) and Training (data collection for models).

From 15 September 2026, for new domains, Training and Agent are blocked by default on pages that display ads, while Search stays allowed. On the same day Cloudflare updated how it handles mixed-use crawlers, such as those from Google, Apple and Microsoft that serve both search and training: blocks now apply to them too. In practice, an overly strict setting can remove a site from classic search results as well.

How to check your site in 10 minutes

  1. Open your robots.txt by adding /robots.txt to your site's address. Look for Disallow lines under the user agents in the table, and for a User-agent: * rule with Disallow: /, often left behind after a redesign.
  2. Check your CDN and firewall. On Cloudflare, look at the AI bot settings in the domain's Security section; on other services, look for rules blocking user agents or IP ranges. Anthropic warns that blocking IP addresses can stop crawlers from reading robots.txt, with unpredictable effects.
  3. Make sure your content is in the HTML. A December 2024 Vercel study found that the major AI crawlers don't run JavaScript: if services, prices and reviews only appear after the page loads, they don't exist for them.
  4. Look at your server logs to see which crawlers actually visit the site and which pages return errors.
  5. Give changes time. OpenAI says it takes about 24 hours for a robots.txt update to reach its systems.

In Geosnap's technical analysis these checks are automatic: robots.txt is read for every AI crawler, and errors that block bots show up among the actions to take. To understand which sites AI relies on when it can't read yours, see where LLMs get their information.

Sources

Frequently asked questions

If I block GPTBot, will my site disappear from ChatGPT?

No. GPTBot collects content to train models. What counts for ChatGPT search answers is OAI-SearchBot: if you let it in, your site stays citable even with GPTBot blocked.

Does blocking Google-Extended hurt my Google rankings?

According to Google, no: Google-Extended does not affect inclusion in Google Search or ranking. It covers the use of content to train Gemini and for grounding.

How long does a robots.txt change take to apply?

OpenAI says about 24 hours for its systems. Other providers don't give precise timings: check again after a few days, including in your server logs.

Is robots.txt enough to block every AI crawler?

No. Crawlers acting on a user's request, such as Perplexity-User, may not follow it. Stricter control needs CDN or firewall rules, used carefully so you don't block search crawlers too.