
AI Crawler Analytics: Which AI Bots Visit Your Site, and How to See Them
Carlos Garcia10/8/2026Every AI assistant that cites your site had to read it first, and it read it with a crawler that identifies itself. Those visits are in your server logs right now, with names like OAI-SearchBot and PerplexityBot and ClaudeBot, and most site owners have never looked at them. The tracking tools call this "agent analytics" and charge for a dashboard of it. You can get the useful part of that dashboard in an afternoon, and the first thing it usually shows is that a robots.txt line written in 2023 is quietly keeping you out of ChatGPT's answers.
This is a guide to which AI crawlers exist, what each one is for, how to see them on your site, and which ones you want to let in. Everything about the crawlers below comes from the vendors' own documentation as of October 2026, linked so you can check it.
The crawlers, and what each one does
The important distinction is between three kinds of visit. A training crawler collects pages to train future models; blocking it keeps your content out of training data and has no effect on whether you appear in answers. A search crawler builds the index the assistant searches when a user asks a question; blocking it removes you from the answers. A user-initiated fetcher visits a page because a user asked about it right now; it generally ignores robots.txt, because the user asked.
OpenAI (ChatGPT)
- OAI-SearchBot is the search crawler. OpenAI's documentation says it "is used to surface websites in search results in ChatGPT's search features" and that "sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers." This is the one that matters for citations.
- GPTBot is the training crawler. Disallowing it "indicates a site's content should not be used in training generative AI foundation models." It does not affect search answers.
- ChatGPT-User is the user-initiated fetcher. OpenAI notes that "because these actions are initiated by a user, robots.txt rules may not apply," and that it "is not used to determine whether content may appear in Search."
- OAI-AdsBot checks pages submitted as ChatGPT ads. Irrelevant unless you advertise there.
OpenAI says a robots.txt change takes about 24 hours to take effect for search. The full list is at OpenAI's bots page.
Anthropic (Claude)
- Claude-SearchBot "navigates the web to improve search result quality for users"; disabling it "prevents our system from indexing your content for search optimization."
- Claude-User "supports Claude AI users": it fetches a page when a user's question needs it. Disabling it "prevents our system from retrieving your content in response to a user query."
- ClaudeBot is the training crawler; disabling it "signals that the site's future materials should be excluded from our AI model training datasets."
Anthropic advises against blocking by IP address, because it can stop the bots from reading your robots.txt at all. Details at Anthropic's crawler article.
Perplexity
- PerplexityBot "is designed to surface and link websites in search results on Perplexity" and "is not used to crawl content for AI foundation models." Perplexity recommends allowing it.
- Perplexity-User fetches pages for a user's question and "generally ignores robots.txt rules," since the user requested the fetch.
Both are documented at Perplexity's bots page, with IP ranges for firewall allowlisting.
Google (AI Overviews, AI Mode, Gemini)
- Googlebot is the crawler for Google Search, and AI Overviews and AI Mode are Google Search. Google's AI features documentation says a page needs only to be "indexed and eligible to be shown in Google Search with a snippet" and that "there are no additional technical requirements." It adds: "You don't need to create new machine readable files, AI text files, or markup to appear in these features."
- Google-Extended is a robots.txt token, not a separate crawler, that controls whether crawled content "may be used for training future generations of Gemini models" and for grounding in Gemini apps and Vertex AI. Google states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." Blocking it does not remove you from AI Overviews.
Others worth knowing
- Bingbot feeds Bing and therefore Microsoft Copilot. If you are visible in Bing, you are eligible in Copilot.
- Applebot indexes for Siri and Spotlight; Applebot-Extended is Apple's separate training opt-out.
- CCBot is Common Crawl, whose open dataset many models train on. Blocking it is a training decision, not a search one.
- meta-externalagent (Meta), Amazonbot (Amazon), Bytespider (ByteDance) are training and product crawlers for their respective companies.
How to see which crawlers visit your site
Analytics tools like Google Analytics and Plausible will not show you any of this. Crawlers do not run JavaScript, so they never fire the tracking script. You need the server side.
- Server or CDN logs. Filter the access log by user agent for the names above. On most hosts this is one command; on a CDN like Cloudflare it is the Security or Analytics section, and Cloudflare shows AI crawlers as their own category. Count visits per crawler per day, and look at which URLs each one requests; the pages OAI-SearchBot fetches most are the pages ChatGPT thinks are worth knowing about.
- Your robots.txt. Read it as the crawlers do, line by line. The common accident is a blanket rule added in 2023 when GPTBot first appeared, which on many sites now also blocks the search crawlers that were introduced later. If you see "User-agent: *" followed by "Disallow: /" for anything, or a block on OAI-SearchBot, PerplexityBot or Claude-SearchBot, you have found why you are not cited.
- Your CDN's bot settings. Several CDNs now block AI crawlers by default on new sites, sometimes including the search crawlers. Check the bot management or AI crawler settings, not only robots.txt.
- Referral traffic, separately. Visits from assistants (people clicking a citation) do show up in analytics, under referrers like chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com and claude.ai. In Plausible, open Sources and filter by those; in GA4, make a channel group for them. Clicks from Google AI Overviews are the exception: Search Console counts them as ordinary Google clicks, with no separate line.
Put the two views together and you have what the agent analytics dashboards sell: which assistants read you, which pages they read, and how many people they send. The difference is that the dashboards add a chart.
Which crawlers to allow
For almost every business the answer is the same. Allow every search crawler and user fetcher: Googlebot, Bingbot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User, Applebot. These are how you get cited, and blocking any of them is opting out of that assistant's answers.
Training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent) are a separate decision. The vendors' documentation is consistent that blocking them does not affect search features, so it is a question of whether you want your content in training data, not of visibility. Publishers with licensing concerns block them; most small businesses have no reason to, and there is a plausible case that being in the training data helps a model know your brand exists.
A minimal robots.txt that gets this right needs no special AI lines at all: allow everything, and add explicit Disallow rules only for the training crawlers you have decided to exclude. If you add an llms.txt file (a plain-text summary of your site for assistants), note that Google has said it does not need one; some other assistants read it and some ignore it, so treat it as a courtesy rather than a requirement. SEO Stuff's own is at seo-stuff.com/llms.txt if you want a template.
A fifteen-minute check
Pull one week of server logs, count visits by the crawler names above, and note which have never visited. Read your robots.txt and your CDN's bot settings and fix anything blocking a search crawler. Add the assistant referrers as a channel in your analytics. Then ask each assistant five questions about your category with web search on and see whether you are cited. If the crawlers visit and you are still not in the answers, the problem is the content, not the access, and that is a different post: see how generative engines choose their sources.
The crawler access check is part of the audit in every SEO Stuff Done-For-You order, and the free SEO + AI search audit covers it too, along with how the assistants currently read and cite your pages.
