Open your server logs and you will find visitors you did not have a few years ago: GPTBot, ClaudeBot, PerplexityBot and a growing list of similar names. Some are collecting text to train future models. Some are building a search index that an assistant queries when answering. Some fetch a single page because a person just asked an assistant about it.
Those three jobs have very different consequences for your brand, and treating them as one thing leads to bad decisions, such as blocking everything and then wondering why you never get cited. This guide sorts the main crawlers into groups, lays out the robots.txt trade-offs, and shows how to log them.
A note before we start: bot names, user-agent strings and policies change. Treat the list below as a starting point and check each provider's current documentation before you act on it.
The three jobs a bot can do
| Type | What it does | Triggered by | Effect of blocking |
|---|---|---|---|
| Training | Collects content to train or improve models | The provider's crawl schedule | Your content is less likely to be included in future training data |
| Search / index | Builds an index used to find and cite sources in answers | The provider's crawl schedule | You may not appear as a cited source in that product's search answers |
| User fetch | Retrieves a specific page on behalf of a user | A person asking an assistant about a URL or topic | The assistant cannot read your page at that moment |
Why this matters: blocking a training bot is a statement about how your content is used for model development. Blocking a search or user-fetch bot can cost you visibility in live answers. They are different decisions.
The main crawlers
OpenAI
- GPTBot: collects content that may be used for training OpenAI's models.
- OAI-SearchBot: indexes pages for ChatGPT's search features, which is where citations in ChatGPT search answers come from. OpenAI documents it as separate from training.
- ChatGPT-User: fetches pages when a user asks ChatGPT to look at a link or when it browses on a user's behalf.
Anthropic
- ClaudeBot: collects content that may be used for training Claude.
- Claude-SearchBot: indexes content to improve search result quality for Claude.
- Claude-User: fetches pages when a Claude user asks something that requires visiting a site.
Perplexity
- PerplexityBot: builds Perplexity's search index. Perplexity describes it as for surfacing and linking sites in results, not for training foundation models.
- Perplexity-User: fetches a page when a user's question triggers a visit.
- Google-Extended: this one is different. It is a robots.txt control token, not a separate crawler you will see in logs. It lets you opt content out of use for training Google's generative models and for grounding in Gemini apps. Google's actual crawling is still done by Googlebot, which means you will not see "Google-Extended" in your access logs.
- Googlebot: AI Overviews and AI Mode draw on Google Search's index. Google states that Google-Extended does not affect inclusion in Search, so blocking it does not remove you from AI Overviews. Controls such as
nosnippetand other preview directives are the levers that affect how your content appears in Search features.
Others you may see
- Applebot-Extended: a robots.txt token that lets you opt out of Apple using content crawled by Applebot for training its models.
- CCBot: Common Crawl's bot. Its dataset is widely used in model training by many parties.
- Bytespider and various other crawlers from other AI companies. Names change frequently.
Some providers state that their user-initiated fetchers may not follow robots.txt in the same way as their scheduled crawlers, because a human asked for the page. Read each provider's documentation for the current position.
Deciding what to allow
There is no universal right answer. Here are the trade-offs.
Option A: allow everything
Simple, and it maximizes your chance of being known to models and cited in live answers. The cost is that your content may be used for training, with no control over how.
Option B: allow search and user-fetch, block training
This is the most common middle path. You keep your eligibility to be cited in live answers while opting out of model training where the provider offers a separate bot. The caveats: you are relying on each provider to honor the opt-out, and blocking training does not remove content already collected.
Option C: block everything
Appropriate for some publishers with licensing concerns or paywalled content. Understand the cost: you reduce the ways assistants can read and cite you, so they will describe you from third-party sources instead. For most brands trying to be discovered, this works against the goal.
A sample robots.txt for Option B
# Training crawlers: opt out
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
# Search and user-fetch bots: allow
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
Test any change before shipping it. A stray Disallow: / under User-agent: * blocks everything, and it is a common cause of "we can't be found" problems.
Things robots.txt does not do
- It is a request, not a lock. Reputable crawlers follow it. Others may not.
- It does not delete data already collected.
- It does not stop a CDN or firewall rule from blocking a bot you wanted to allow. Check both layers.
Verifying a bot is real
User-agent strings are trivial to fake. Anyone can send a request that claims to be GPTBot. Before you draw conclusions or write firewall rules:
- Check whether the provider publishes IP ranges for its bots, and compare the request's IP against them.
- Use reverse DNS verification where the provider supports it.
- Treat unverified hits as "claims to be," and keep them separate in your analysis.
How to log AI crawler visits
Standard web analytics tools such as JavaScript trackers will not see crawlers, because most bots do not execute your page scripts. You need server-side or edge-side logging.
Option 1: CDN worker or edge function
If you use a CDN with edge compute, run a small function on each request that checks the user-agent against a list of known AI bots and, on a match, sends a log event: bot name, user-agent, path, status code, IP and timestamp. Fire the log call asynchronously so it does not slow the response.
The shape of the logic:
const BOTS = [
'GPTBot',
'OAI-SearchBot',
'ChatGPT-User',
'ClaudeBot',
'Claude-SearchBot',
'Claude-User',
'PerplexityBot',
'Perplexity-User',
];
export default {
async fetch(request, env, ctx) {
const response = await fetch(request);
const ua = request.headers.get('user-agent') || '';
const bot = BOTS.find((b) => ua.includes(b));
if (bot) {
ctx.waitUntil(
fetch(env.LOG_ENDPOINT, {
method: 'POST',
headers: {
'content-type': 'application/json',
authorization: `Bearer ${env.LOG_TOKEN}`,
},
body: JSON.stringify({
bot,
user_agent: ua,
path: new URL(request.url).pathname,
status_code: response.status,
ip: request.headers.get('cf-connecting-ip'),
visited_at: new Date().toISOString(),
}),
}),
);
}
return response;
},
};
Option 2: application middleware
If you do not use edge compute, add middleware in your web framework that does the same check after the response is built and queues a log event. This catches only requests that reach your application, so cached pages served by a CDN will be missed.
Option 3: parse server logs
Ship your web server's access logs somewhere you can query and filter by user-agent. It is the least intrusive approach, but it is batch and typically needs more setup to turn into readable reports.
What to look at once you have logs
- Which bots visit, and how often. A bot that never appears may be blocked by you, or may simply not have discovered you yet.
- Which paths they fetch. Do they reach your key pages, or only your homepage?
- Status codes. Repeated 403 or 429 responses suggest a firewall or rate limit is turning them away. Repeated 404s point at stale links.
- Training vs search vs user-fetch mix. A burst of user-fetch hits on a page can mean people are asking assistants about it.
This is where crawler data connects to the rest of your measurement. If a page is cited in answers, you should expect search or user-fetch bots to have read it. If a page you care about is never fetched, fix access before you rewrite the content. For how the pieces fit together, see how AI assistants choose brands.
How Citeflare helps
Citeflare gives each project an ingest endpoint and a ready-made Cloudflare Worker snippet, so you can start logging AI crawler visits without building your own pipeline. Hits are classified by bot, and you can see which pages each crawler reads and what status it received. Explore crawler analytics or create an account to connect your site.