# Is your website blocking ChatGPT without you knowing? A guide to AI crawlers

> GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot: what each AI crawler does, how to set up robots.txt and what Cloudflare's new rules change for your site.

URL: https://geosnap.ai/en/blog/il-tuo-sito-blocca-chatgpt-guida-crawler-ai
Language: English
Version IT: https://geosnap.ai/blog/il-tuo-sito-blocca-chatgpt-guida-crawler-ai
Author: Rinald Sefa, CMO Geosnap · Category: Guides
Publisher: Geosnap (Maind Group S.r.l.), https://geosnap.ai

## In short

ChatGPT, Claude and Perplexity can only cite you if their search crawlers can read your site. You can block model training and stay citable, but you need to tell the right crawlers apart in robots.txt and check your CDN and firewall too, starting with Cloudflare's new settings of 15 September 2026.

- Every AI provider uses different crawlers for search, for training and for user requests.
- Blocking GPTBot or ClaudeBot stops training; blocking OAI-SearchBot or Claude-SearchBot removes you from answers.
- Since 15 September 2026 Cloudflare blocks Training and Agent by default on ad pages of new domains.
- The check takes ten minutes: robots.txt, CDN settings, content in the HTML and server logs.

Written by the Geosnap AI agent, reviewed and approved by Rinald Sefa.

Translated from the Italian original. [Read the original](https://geosnap.ai/blog/il-tuo-sito-blocca-chatgpt-guida-crawler-ai)

If ChatGPT, Claude or Perplexity never cite your website, the cause may be more mundane than you think: a rule in your `robots.txt` file, a CDN setting or a firewall that keeps their crawlers out. Nobody notices, because people still see the site as usual. This guide covers which AI crawlers exist, what each one does, how to decide which ones to let in, and what changed with Cloudflare's new settings.

## Why robots.txt matters for getting cited by AI

AI assistants can only cite pages they can read. Before answering a question, ChatGPT, Claude and Perplexity search the web with their own crawlers: if your site blocks them, someone else shows up in the answer.

`robots.txt` is the file at the root of your site that tells crawlers what they may visit. The AI crawlers of the major providers say they respect it, so a single `Disallow` line is enough to make you disappear from their searches. The problem is that many of these lines were added in recent years to block model training, without distinguishing between crawlers that collect training data and crawlers looking for sources to cite.

## Which AI crawlers exist and what they do

Each provider runs several crawlers with different purposes. These are the officially documented ones:

| Provider | User agent | Purpose | Respects robots.txt |
| --- | --- | --- | --- |
| OpenAI | `OAI-SearchBot` | Finding sites to show in ChatGPT search | Yes |
| OpenAI | `GPTBot` | Collecting content to train models | Yes |
| OpenAI | `ChatGPT-User` | Visiting a page when a user asks | Not always |
| Anthropic | `Claude-SearchBot` | Indexing content for Claude's searches | Yes |
| Anthropic | `ClaudeBot` | Collecting content for training | Yes |
| Anthropic | `Claude-User` | Reading a page at a user's request | Yes |
| Perplexity | `PerplexityBot` | Surfacing and linking sites in Perplexity results | Yes |
| Perplexity | `Perplexity-User` | Visiting a page at a user's request | Generally no |
| Google | `Google-Extended` | Use of content to train Gemini and for grounding | Yes |

OpenAI explains that [sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers](https://developers.openai.com/api/docs/bots), at most as navigational links. Perplexity notes that [Perplexity-User generally ignores robots.txt](https://docs.perplexity.ai/guides/bots), because a person requested the visit. Google, for its part, states that [Google-Extended does not affect inclusion in Google Search or ranking](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers).

## Search and training can be separated

Yes: you can keep your content out of model training and still be citable. OpenAI says so explicitly: [each setting is independent of the others](https://developers.openai.com/api/docs/bots), so you can allow OAI-SearchBot and block GPTBot. The same goes for Anthropic's crawlers: according to [Claude's documentation](https://support.claude.com/en/articles/8896518-what-is-claudebot), blocking ClaudeBot excludes future content from training, while blocking Claude-SearchBot may reduce your visibility in results.

A `robots.txt` that stays visible in answers without contributing to training can look like this:

```
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /
```

It's a choice, not a rule: many brands prefer to leave training open too, because models learn about the brand. What matters is that it's a deliberate decision, not a side effect.

## Cloudflare: what changed on 15 September 2026

If your site runs through Cloudflare, check its bot settings: CDN rules apply even when robots.txt is open. In July 2026 Cloudflare [split AI crawlers into three categories](https://blog.cloudflare.com/content-independence-day-ai-options/): Search (indexing for search results), Agent (automated actions on behalf of users) and Training (data collection for models).

From 15 September 2026, for new domains, Training and Agent are blocked by default on pages that display ads, while Search stays allowed. On the same day Cloudflare [updated how it handles mixed-use crawlers](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/), such as those from Google, Apple and Microsoft that serve both search and training: blocks now apply to them too. In practice, an overly strict setting can remove a site from classic search results as well.

## How to check your site in 10 minutes

1. **Open your robots.txt** by adding `/robots.txt` to your site's address. Look for `Disallow` lines under the user agents in the table, and for a `User-agent: *` rule with `Disallow: /`, often left behind after a redesign.
2. **Check your CDN and firewall.** On Cloudflare, look at the AI bot settings in the domain's Security section; on other services, look for rules blocking user agents or IP ranges. Anthropic warns that [blocking IP addresses can stop crawlers from reading robots.txt](https://support.claude.com/en/articles/8896518-what-is-claudebot), with unpredictable effects.
3. **Make sure your content is in the HTML.** A [December 2024 Vercel study](https://vercel.com/blog/the-rise-of-the-ai-crawler) found that the major AI crawlers don't run JavaScript: if services, prices and reviews only appear after the page loads, they don't exist for them.
4. **Look at your server logs** to see which crawlers actually visit the site and which pages return errors.
5. **Give changes time.** OpenAI says it takes [about 24 hours](https://developers.openai.com/api/docs/bots) for a robots.txt update to reach its systems.

In [Geosnap's technical analysis](https://geosnap.ai/en/features) these checks are automatic: robots.txt is read for every AI crawler, and errors that block bots show up among the actions to take. To understand which sites AI relies on when it can't read yours, see [where LLMs get their information](https://geosnap.ai/en/blog/sources-ai-da-dove-prendono-informazioni-llm).

## Sources

- [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) (OpenAI, 2026)
- [Does Anthropic crawl data from the web, and how can site owners block the crawler?](https://support.claude.com/en/articles/8896518-what-is-claudebot) (Anthropic, 2026)
- [Perplexity Crawlers](https://docs.perplexity.ai/guides/bots) (Perplexity, 2026)
- [Google's common crawlers](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) (Google Search Central, 2026)
- [Your site, your rules: new AI traffic options for all customers](https://blog.cloudflare.com/content-independence-day-ai-options/) (Cloudflare, 2026-07-01)
- [Have it both ways: stay discoverable in search while disallowing AI training](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/) (Cloudflare, 2026-09-15)
- [The rise of the AI crawler](https://vercel.com/blog/the-rise-of-the-ai-crawler) (Vercel, 2024-12-17)

## Frequently asked questions

### If I block GPTBot, will my site disappear from ChatGPT?

No. GPTBot collects content to train models. What counts for ChatGPT search answers is OAI-SearchBot: if you let it in, your site stays citable even with GPTBot blocked.

### Does blocking Google-Extended hurt my Google rankings?

According to Google, no: Google-Extended does not affect inclusion in Google Search or ranking. It covers the use of content to train Gemini and for grounding.

### How long does a robots.txt change take to apply?

OpenAI says about 24 hours for its systems. Other providers don't give precise timings: check again after a few days, including in your server logs.

### Is robots.txt enough to block every AI crawler?

No. Crawlers acting on a user's request, such as Perplexity-User, may not follow it. Stricter control needs CDN or firewall rules, used carefully so you don't block search crawlers too.
