All resources
Blog

How to Read Your Server Logs for AI Crawlers: A Checklist

Which AI bots show up in your access logs, what each one does (training, search indexing or user-triggered fetch), how to verify they are real, and a checklist of what to fix.

Your server keeps a diary. Every time something asks for a page, the server writes one line saying who asked, when, for which page, and what the server answered. This diary is called an access log.

AI companies now show up in that diary every day. Some of their bots collect pages to train AI models. Some build a search index that AI answers draw from. Some fetch a page only because a real person just asked a chatbot a question. These three jobs mean very different things for your business, and your logs can tell them apart.

Why AI bots in your logs are worth a look

A "user agent" is the name tag a visitor sends with each request. A browser sends one. So does every well-behaved bot. When OpenAI's search crawler visits, for example, its name tag contains the word OAI-SearchBot.

The amount of AI bot traffic is no longer small. Cloudflare, a network company that handles traffic for many websites, has published these figures:

FindingPeriodSource
80% of AI crawling was for training, 18% for search and 2% for user actions12 months to August 2025Cloudflare1
Anthropic crawled about 38,000 pages for every visit it sent back; OpenAI about 1,091; Perplexity about 194July 2025Cloudflare1
AI "user action" crawling grew by more than 15x2025Cloudflare Radar2
Other AI bots made up 4.2% of HTML requests, while Googlebot alone made up 4.5%2025Cloudflare Radar2
52% of crawler requests were for AI training, up from 22% in spring 2025June 2026Cloudflare3

The second row needs a quick explanation. Cloudflare calls it the "crawl-to-refer ratio." It compares how many pages a company's bots fetch with how many visitors that company sends back to websites.1 A high number means lots of reading and few clicks in return.

The bots that may bring readers back (search and user-triggered fetchers) are not the ones that only take pages for training. Mixing them up leads to bad decisions, like blocking the bot that would have cited you.

The three jobs AI bots do

Almost every AI user agent fits one of three jobs.

  1. Training. The bot collects public pages that may be used to train future AI models. Blocking it is how you opt out of that training. It does not decide whether you appear in today's AI answers.
  2. Search indexing. The bot builds a search index. AI answer engines pull from that index when they write answers and pick links. If you block it, you may disappear from that company's AI search results.
  3. User-triggered fetch. A real person asked a chatbot something, and the chatbot went to fetch a specific page to answer. These are not bulk crawls. Each hit usually traces back to one live question.

There is also a fourth thing that looks like a bot but isn't: a control token. Google-Extended and Applebot-Extended are names you can use in robots.txt (the text file at the root of your site that tells bots what they may and may not crawl). But they never appear in your logs. They only tell the company how it may use pages its normal crawler already fetched.

Reference table: the AI user agents that matter

The table below comes from each company's own documentation. The "robots.txt" column states only what each company's page says.

CompanyUser agentJobRespects robots.txt? (per its docs)How to verify
OpenAIGPTBotTrainingYes. Disallowing it is the opt-out from training.4IP list at openai.com/gptbot.json4
OpenAIOAI-SearchBotSearch index for ChatGPT search featuresYes. Sites that opt out won't show in ChatGPT search answers, though they can still appear as navigational links.4IP list at openai.com/searchbot.json4
OpenAIChatGPT-UserUser-triggered fetch (ChatGPT and Custom GPTs)"Because these actions are initiated by a user, robots.txt rules may not apply."4IP list at openai.com/chatgpt-user.json4
AnthropicClaudeBotTrainingYes. Anthropic says its bots honor robots.txt and gives a Crawl-delay example for ClaudeBot.5IP list at claude.com/crawling/bots.json5
AnthropicClaude-SearchBotSearch quality and indexingYes (covered by the same statement about Anthropic's bots)5Same Anthropic IP list5
AnthropicClaude-UserUser-triggered fetchYes (covered by the same statement). Anthropic warns that blocking it may reduce visibility in user-directed web search.5Same Anthropic IP list5
PerplexityPerplexityBotSearch index. Perplexity says it is not used to crawl content for AI foundation models.6Perplexity recommends allowing it in robots.txt so your site appears in results.6IP list at perplexity.com/perplexitybot.json6
PerplexityPerplexity-UserUser-triggered fetchGenerally ignores robots.txt, because a user requested the fetch.6IP list at perplexity.com/perplexity-user.json6
GoogleGooglebotSearch index, including AI Overviews and AI ModeYes. robots.txt rules for Googlebot are the crawl control for Search.7Reverse DNS to googlebot.com, google.com or googleusercontent.com, or Google's IP lists8
GoogleGoogle-ExtendedControl token only (Gemini training and grounding)It is a robots.txt token. It has no separate user agent in requests.9Nothing to verify. It never appears in logs.
GoogleGoogle-Agent and other user-triggered fetchersUser-triggered fetch"Generally ignore robots.txt rules."10Reverse DNS to gae.googleusercontent.com or google-proxy hosts, or Google's user-triggered IP lists8
MicrosoftBingbotBing search indexBing says it respects preferences set in robots.txt.11Reverse DNS should end in search.msn.com, then forward DNS back to the same IP12
AppleApplebotSearch (Spotlight, Siri, Safari), and may help train Apple's modelsYes. Follows rules for Applebot, and falls back to Googlebot rules if none exist. Ignores crawl-delay.13Reverse DNS to applebot.apple.com, or the IP list at search.developer.apple.com/applebot.json13
AppleApplebot-ExtendedControl token only (training opt-out)Apple says it "does not crawl webpages."13Nothing to verify. It never appears in logs.
Metameta-externalagentTraining, or improving products by indexing contentMeta says to block its crawlers with a disallow line in robots.txt and mentions no bypass for this one.14Compare the IP with Meta's route list for AS32934, fetched with a whois query14
Metameta-webindexerSearch index for Meta AIMeta says allowing it in robots.txt helps Meta cite and link your content in Meta AI's responses.14Same Meta route list14
Metameta-externalfetcherUser-triggered fetch"This crawler may bypass robots.txt rules."14Same Meta route list14
Common CrawlCCBotOpen repository of web crawl data that anyone can useCommon Crawl tells you to block it with a CCBot group in robots.txt.15Reverse DNS to crawl.commoncrawl.org, or the IP list at index.commoncrawl.org/ccbot.json15

A few details from the table are easy to miss:

  • User-triggered fetchers don't follow one rule. OpenAI says robots.txt "may not apply." Perplexity and Google say their fetchers generally ignore it. Meta says its fetcher may bypass it. Anthropic's statement that its bots honor robots.txt covers Claude-User too. So a block line in robots.txt can mean different things for each company.4561014
  • Google-Extended does not touch Search. Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."9 AI Overviews and AI Mode use Googlebot and normal Search controls. A page must be indexed and eligible for a snippet to appear there.7
  • robots.txt changes are not instant. OpenAI says it can take about 24 hours for its search systems to adjust.4 Meta says crawlers may cache robots.txt for up to 24 hours.14
  • Version numbers change. OpenAI's GPTBot currently identifies as GPTBot/1.4, and OpenAI notes the version may change.4 Always search your logs for the bot name, not the full string.

How to check a bot is really who it says it is

Anyone can type "GPTBot" into a user agent. Common Crawl says plainly that it is aware of crawlers falsely calling themselves CCBot.15 Google warns that user agent strings can be spoofed too.10 So a name in your log is a claim, not proof.

There are two ways to check the claim. Both use the visitor's IP address, which is the first thing on each log line.

Method 1: match the IP against the company's published list

OpenAI, Anthropic, Perplexity, Google, Apple and Common Crawl all publish JSON files listing the IP ranges their bots use. You check whether the visitor's IP falls inside one of those ranges. If it does, the visit is genuine. If it doesn't, treat it as an impostor.

These lists change. At the time of writing, Anthropic's file showed a creation date of August 18, 2026, and OpenAI's GPTBot file showed September 22, 2026. Download a fresh copy each time instead of saving one forever.

Method 2: reverse DNS, then forward DNS

DNS is the internet's phone book. A reverse lookup asks "what name belongs to this IP?" A forward lookup asks "what IP belongs to this name?" Google's steps go like this:8

  1. Run a reverse lookup on the IP from your log, using the host command.
  2. Check that the name ends in googlebot.com, google.com or googleusercontent.com.
  3. Run a forward lookup on that name.
  4. Check that it gives back the same IP you started with.

Google's own example uses host 66.249.66.1, which returns a name like crawl-66-249-66-1.googlebot.com. You then run host crawl-66-249-66-1.googlebot.com and confirm it points back to 66.249.66.1.8 The same two-step check works for Bingbot (names ending in search.msn.com)12, Applebot (applebot.apple.com)13 and CCBot (crawl.commoncrawl.org)15.

What an unverified "AI bot" can mean

Most fakes are scrapers borrowing a famous name. The reverse can also happen. In August 2025, Cloudflare reported that when Perplexity's declared crawlers were blocked, it saw requests using a generic Chrome user agent and IPs outside Perplexity's published ranges.16 Perplexity disputed this, calling Cloudflare's post a "sales pitch."17 In the same test, Cloudflare said ChatGPT-User fetched robots.txt and stopped when it was disallowed.16

Worked example: classifying six log lines

These lines are illustrative. We wrote them for this guide. The IP addresses come from ranges reserved for documentation (203.0.113.x, 198.51.100.x and 192.0.2.x), so they will never verify against any real bot list. The site is a made-up shop, example-outdoor.com. The format is the common "combined" log format used by Apache and Nginx. The user agent strings follow the formats the companies publish.

  1. 203.0.113.10 - - [07/Oct/2026:08:14:02 +0000] "GET /tents/ultralight-2p HTTP/1.1" 200 48211 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"
  2. 203.0.113.25 - - [07/Oct/2026:08:15:40 +0000] "GET /guides/how-to-choose-a-sleeping-bag HTTP/1.1" 200 61034 "-" "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot"
  3. 198.51.100.7 - - [07/Oct/2026:09:02:11 +0000] "GET /tents/ultralight-2p HTTP/1.1" 200 48211 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"
  4. 198.51.100.44 - - [07/Oct/2026:09:20:57 +0000] "GET /returns-policy HTTP/1.1" 404 512 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)"
  5. 192.0.2.80 - - [07/Oct/2026:10:05:33 +0000] "GET /sale HTTP/1.1" 301 0 "-" "Claude-User"
  6. 192.0.2.99 - - [07/Oct/2026:11:41:09 +0000] "GET /wp-login.php HTTP/1.1" 404 512 "-" "GPTBot"

Each line reads left to right: IP, time, the page asked for, the status code the server sent back, the size, and the user agent. A status code is a three-digit answer. 200 means "here's the page." 301 means "it moved, go here instead." 404 means "not found."

LineClaimed botJobWhat it tells youAction
1GPTBotTrainingA product page was collected for possible model training. It says nothing about appearing in ChatGPT answers today.Verify the IP against gptbot.json. Then decide on a training policy as a business choice.
2OAI-SearchBotSearch indexYour buying guide is being considered for ChatGPT search features. Good news if verified.Verify against searchbot.json. Make sure the page loads fast and the key answer is in the HTML.
3ChatGPT-UserUser-triggered fetchSomeone asked ChatGPT a question at about 09:02 that led it to this exact tent page. That's a live signal of real demand.Log it as a "prompt signal." Check the page answers likely questions (weight, price, size) clearly.
4PerplexityBotSearch indexAn AI search bot asked for your returns page and got a 404. It may be following an old link.Fix fast. Redirect /returns-policy to the live page, or update the link that points there.
5Claude-UserUser-triggered fetchA Claude user's question led to /sale, which redirects. Anthropic's help page doesn't publish a full user agent string, so the name alone proves little.Check the IP against Anthropic's bots.json. If genuine, make sure the redirect is a single hop to a live page.
6"GPTBot"Probably fakeA login page probe with a bare "GPTBot" name. Real training crawlers don't hunt for WordPress login pages.Verify the IP. If it isn't in OpenAI's list, treat it as a scraper and block it by IP at your firewall or CDN, not in robots.txt.

Lines 1 and 3 hit the same page. One is a training crawl. The other is a person asking a question right now. That's why sorting by job matters more than counting total "AI hits."

Commands to count AI bot hits per bot and per URL

These commands work on Mac or Linux, on a standard combined-format log file named access.log. Ask your developer or host for the raw access log covering the last 30 days. If your log uses a different format, the column numbers in the awk commands may need to change.

1. Count hits for each AI bot.

for b in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot Claude-User PerplexityBot Perplexity-User Googlebot bingbot Applebot meta-externalagent meta-webindexer meta-externalfetcher CCBot; do printf "%s\t%s\n" "$b" "$(grep -ci "$b" access.log)"; done

2. See which pages one bot asked for most. In the combined format, the page path is column 7.

grep "OAI-SearchBot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

3. See which status codes a bot received. The status code is column 9.

grep "OAI-SearchBot" access.log | awk '{print $9}' | sort | uniq -c | sort -rn

4. List the pages where AI search and user bots hit errors or redirects.

grep -E "OAI-SearchBot|ChatGPT-User|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User" access.log | awk '$9 ~ /^(301|302|404|410|5)/ {print $9, $7}' | sort | uniq -c | sort -rn | head -30

5. Pull the unique IPs claiming to be one bot, ready for checking.

grep "GPTBot" access.log | awk '{print $1}' | sort -u > gptbot_ips.txt

6. Check those IPs against the official list (needs Python 3, which most Macs and Linux machines have).

python3 -c "import json,ipaddress,urllib.request; nets=[ipaddress.ip_network(p.get('ipv4Prefix') or p.get('ipv6Prefix')) for p in json.load(urllib.request.urlopen('https://openai.com/gptbot.json'))['prefixes']]; [print(ip, 'OK' if any(ipaddress.ip_address(ip) in n for n in nets) else 'NOT IN LIST') for ip in open('gptbot_ips.txt').read().split()]"

Swap the URL for searchbot.json, chatgpt-user.json, Anthropic's bots.json or Perplexity's files to check the other bots. Their files use the same "prefixes" structure.

The spreadsheet method (no command line)

If you don't use a terminal, a spreadsheet does the same job for a small site.

  1. Get the access log from your host as a text file. Open it in Google Sheets or Excel.
  2. Put each full log line in column A. Don't worry about splitting it yet.
  3. In column B, add the bot names from the reference table, one per row.
  4. In column C, count each bot with a formula like =COUNTIF(A:A,"*"&B2&"*"). Drag it down.
  5. To see one bot's pages, use the filter tool on column A with "contains: OAI-SearchBot." Then scan the page paths.
  6. To find errors, add a second filter for " 404 " (with spaces on both sides) or " 301 ".
  7. Add a column called "Job" and label each bot as Training, Search or User. Then make a pivot table of hits by job and by week.

This won't verify IPs. For that, copy the unique IPs to your developer, or use the command in step 6.

Two traps that hide real traffic

  • Your CDN may answer first. If you use a CDN (a network that stores copies of your pages closer to visitors), cached pages may never reach your server. Your server log then undercounts bots. Check your CDN's own logs too.
  • The IP may be your CDN's, not the bot's. Behind a proxy, column 1 may show the proxy's IP. The real visitor IP is often in a separate header field. Ask your developer which field holds it before you verify anything.

The checklist: what to act on

Once you can sort hits by bot and by job, work through this list.

Access and robots.txt

  • Confirm you're not blocking AI search bots by accident. A blanket "block AI" rule can catch more than training bots. A rule that blocks OAI-SearchBot, Claude-SearchBot or PerplexityBot can remove you from those companies' AI search results.456
  • Decide training and search separately. You can block GPTBot and still allow OAI-SearchBot.4 You can disallow Google-Extended and stay in Google Search.9 Write down your policy for each job.
  • Check for wildcard blocks. A User-agent: * group with broad Disallow rules applies to every bot that has no group of its own. Applebot also falls back to your Googlebot rules if it has none.13
  • Don't rely on robots.txt for user-triggered fetchers. Several companies say these may skip it.461014 If a page must stay private, put it behind a login.
  • Check your firewall and CDN bot rules. A security setting can block verified AI search bots even when robots.txt allows them. Look for a jump in 403 (forbidden) codes for those bots.

Errors and wasted crawls

  • Fix 404s that AI search and user bots keep hitting. A bot asking for a missing page usually means a link somewhere still points to it. Redirect it to the right live page.
  • Shorten redirect chains. One hop (old URL to new URL) is fine. Two or three hops waste the bot's time and add failure points.
  • Watch for 5xx errors. These mean your server failed. If bots get many of them, they may slow down or come back later.
  • If bots overload your server, use the right signal. Google already reads 500, 503 and 429 responses as a sign to slow down. In October 2026 it added examples showing that a Retry-After header on a 503 or 429 tells its crawlers when to try again.18 Use this for emergencies, not as a daily setting.

Coverage

  • List important pages AI search bots never fetched. Compare your top commercial pages with the URL list from command 2. A key page with zero visits from any search bot may be poorly linked or missing from your sitemap.
  • Give new pages time before you worry. Search Engine Roundtable reported on October 5, 2026 that Google's Gary Illyes showed internal data at a Search Central Live event: new URL discovery typically takes about 20 hours, but can take weeks or never happen.19
  • Check that key facts are in the raw HTML. If prices, specs or answers only load through JavaScript, a bot may fetch the page and still miss them. View the page source and look for the text.
  • For Googlebot, use the Crawl Stats report. In Search Console, under Settings, it shows requests, response codes, file types and Googlebot type over time. It only works for root-level properties.20

Signals of real demand

  • Track user-triggered fetches by page and by week. ChatGPT-User, Claude-User and Perplexity-User hits are the closest thing in your logs to "someone just asked an AI about this." Rising hits on a page suggest rising interest in its topic.
  • Read the timing. Cloudflare found user-action traffic follows a clear daily cycle, unlike the more erratic pattern of training crawls.21 Bursts on specific product or comparison pages are worth noting for your sales and content teams.
  • Treat these pages as high priority. If real prompts lead to a page, it should answer clearly, load fast and return a 200.

Hygiene

  • Verify before you block. Blocking the wrong IP range can shut out a genuine bot.
  • Repeat monthly. New user agents appear. OpenAI's page, for example, now also lists OAI-AdsBot, which checks landing pages submitted as ChatGPT ads.4

What logs can't tell you

A log hit is not a citation. A citation is when an AI answer names or links your page as a source. Your log only shows that a bot fetched a page. It doesn't show whether the page was used, quoted, linked or ignored.

Here's why:

  • A training crawl can feed a model that never mentions you.
  • A search bot can index a page that never wins a spot in an answer.
  • A user-triggered fetch shows a page was read for one question. The final answer might still recommend a competitor.
  • Google says robots.txt rules for Googlebot are the crawl control for its AI features in Search, just as for regular results.7 So your log can't separate "fetched for blue links" from "fetched for an AI answer."

Some platforms now report citations directly. Bing Webmaster Tools added an AI Performance report in February 2026. It shows how often your pages are cited in Microsoft Copilot, AI summaries in Bing and some partner tools. Even Bing says these numbers don't show ranking or placement within an answer.11

So use logs for what they're good at: access, errors, coverage and demand signals. Then measure visibility separately, by checking what AI answers actually say about you.

See what AI answers say once the bots have read your pages

Your logs can't show whether ChatGPT, Perplexity, Gemini or Google's AI features actually cite you, or what they say when they do. Ripplix's free report checks the answers themselves, so you can connect "the bot fetched it" to "the AI recommended it."

Get your free AI Visibility Report →

Sources:

  1. Cloudflare: The crawl-to-click gap: Cloudflare data on AI bots, training, and referrals (August 29, 2025)
  2. Cloudflare: The 2025 Cloudflare Radar Year in Review (December 15, 2025)
  3. Cloudflare: Content Independence Day, one year on: building the business model for the agentic Internet (July 1, 2026)
  4. OpenAI: Overview of OpenAI Crawlers (accessed October 8, 2026)
  5. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated April 7, 2026)
  6. Perplexity: Perplexity Crawlers (accessed October 8, 2026)
  7. Google Search Central: AI features and your website (updated December 10, 2025)
  8. Google Search Central: Verifying Googlebot and other Google crawlers (accessed October 8, 2026)
  9. Google Search Central: Google's common crawlers (accessed October 8, 2026)
  10. Google Search Central: Google's user-triggered fetchers (accessed October 8, 2026)
  11. Bing Webmaster Blog: Introducing AI Performance in Bing Webmaster Tools Public Preview (February 10, 2026)
  12. Bing Webmaster Blog: How to Verify that Bingbot is Bingbot (August 31, 2012)
  13. Apple: About Applebot (accessed October 8, 2026)
  14. Meta for Developers: Meta Web Crawlers (accessed October 8, 2026)
  15. Common Crawl: CCBot (accessed October 8, 2026)
  16. Cloudflare: Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives (August 4, 2025)
  17. TechRepublic: Cloudflare Accuses AI Startup of 'Stealth Crawling Behavior' Across Millions of Sites (August 5, 2025)
  18. Search Engine Roundtable: Google crawl rate documentation update (October 6, 2026)
  19. Search Engine Roundtable: Google Search Data On Crawling, Indexing & Serving Timelines (October 5, 2026)
  20. Google Search Console Help: Crawl Stats report (accessed October 8, 2026)
  21. Cloudflare: A deeper look at AI crawlers: breaking down traffic by purpose and industry (August 28, 2025)