This report is updated monthly with fresh Cloudflare Radar data. Bookmark this page to track how AI crawlers are reshaping web traffic each month.
Last month the headline was ByteDance's Bytespider surging to #4 while Applebot's one-month spike unwound. This month the pattern repeated one rung higher — and with a new protagonist. Bytespider's own surge reversed (10.1% → 7.3%), sliding back to #6, exactly the way Applebot faded a month earlier. Taking its place near the top is the crawler that has quietly V-shaped all year: Anthropic's ClaudeBot exploded +66% to 20.0%, leapfrogging Meta-ExternalAgent and GPTBot to become the second-largest AI crawler on the web — behind only Googlebot. It's the single largest one-month share gain any crawler has posted in the life of this report. Meanwhile Googlebot kept sliding (26.7% → 24.9%, below 25% for the first time), and last quarter's "Training will exceed 55% by June" call finished the quarter decisively wrong: training crawling sits at 47.4% while Search crawling hit a record 10.5%, clearing the 10% line I forecast in the May edition.
I analyzed the latest 28 days of Cloudflare Radar AI Insights data — the window covering June 8 through July 6, 2026 — and compared it against the preceding 28-day window (May 11 through June 8) so every month-over-month figure below is computed on an identical basis. The picture is the same diversification story I've tracked all year, but with a twist: for the first time in 2026, top-five concentration rose — not because the incumbents held the line, but because one of them (Anthropic) scaled up faster than everyone else combined. And the recurring lesson of this report got another data point: whichever crawler spikes hardest one month (Applebot, then Bytespider) tends to give most of it back the next. New this edition, I've also added a section translating the data into concrete E-E-A-T guidance for site owners who want AI systems to cite them. If you want the foundational primer on the pipelines behind all this, our explainer on how search engines really work walks through crawl budgets, inverted indexes, and learning-to-rank — all the systems these bots are feeding.
📊 Stats Alert: ClaudeBot grew +66% month-over-month (12.1% → 20.0%) to become the #2 AI crawler, passing Meta-ExternalAgent and GPTBot. Bytespider's surge reversed, falling from 10.1% to 7.3% (#6). Googlebot dropped below 25% to 24.9% — a new low. Training crawling climbed to 47.4% while Search crawling hit a record 10.5%, led by Claude-SearchBot rising to 3.3%. On Workers AI, Text Embeddings surged to 82.4% of inferences as edge workloads pivoted from speech to retrieval.
Who Are the Top AI Crawlers in June 2026?
Googlebot remains the largest AI-related crawler globally, but its decline continued — from 26.7% to 24.9%, a -1.8 pp drop that takes it below 25% for the first time in Cloudflare Radar's tracking. The story just behind it is a full reshuffle: ClaudeBot rocketed from 12.1% to 20.0% (#2), passing both Meta-ExternalAgent and GPTBot in a single month. Meta-ExternalAgent fell to 10.2% (#3, third straight monthly decline), GPTBot eased to 9.6% (#4), and Bytespider's surge reversed hard, dropping -2.8 pp to 7.3% and surrendering the #4 spot it had just taken — sliding all the way to #6, behind Bingbot.
The other headline is again in the search tier: Claude-SearchBot rose to 3.3% (+0.6 pp), extending its lead as the single largest dedicated AI search crawler on the web. Anthropic now leads on both fronts — the fastest-growing training crawler (ClaudeBot) and the largest search crawler (Claude web search) — a level of end-to-end crawling scale no other operator matched this month.
| AI Bot | June 2026 Share (%) | May 2026 Share (%) | Operator | Primary Purpose |
|---|---|---|---|---|
| Googlebot | 24.9% | 26.7% | Search indexing + AI training (mixed) | |
| ClaudeBot | 20.0% | 12.1% | Anthropic | Model training |
| Meta-ExternalAgent | 10.2% | 12.6% | Meta | AI training |
| GPTBot | 9.6% | 10.1% | OpenAI | Model training |
| Bingbot | 8.1% | 8.2% | Microsoft | Search indexing + AI (mixed) |
| Bytespider | 7.3% | 10.1% | ByteDance | AI training |
| Applebot | 5.8% | 6.9% | Apple | Search + AI features |
| Amazonbot | 5.5% | 5.4% | Amazon | AI training |
| Claude-SearchBot | 3.3% | 2.7% | Anthropic | Claude web search |
Here's the twist on the year-long diversification story: top-five company concentration — Google, Meta, OpenAI, Anthropic, and Microsoft, counting each company's primary crawler — rose to 72.8% in June, up from 69.6% in May. That's the first monthly increase in concentration all year, reversing a five-month slide that ran from 84.5% in January. But it isn't the incumbents collectively holding the line — it's Anthropic alone. Between ClaudeBot (20.0%) and Claude-SearchBot (3.3%), Anthropic now accounts for 23.3% of all identified AI crawler traffic, second only to Google and more than Meta's and OpenAI's primary crawlers (10.2% + 9.6%) combined. The market didn't re-consolidate; one challenger scaled up so fast it single-handedly bent the concentration line.
If you're managing AI bot access on your site, the operator that most needs a fresh look this month is Anthropic — ClaudeBot's traffic has nearly doubled in 30 days, and its search crawler is the one most likely to actually cite you. For context on how crawling translates into actual referral traffic back to websites, check our companion Search Engine Referral Report on crawl-to-refer ratios — and see the crawl-to-refer breakdown further down this report.
Which AI Crawlers Gained or Lost Ground This Month?
The headline shift in June is a single-crawler surge offsetting broad declines at the top — the same shape as May, but bigger and driven by a different bot. ClaudeBot (+8.0 pp) absorbed almost all the ground given up by Bytespider (-2.8 pp), Meta-ExternalAgent (-2.4 pp), Googlebot (-1.8 pp), Applebot (-1.0 pp), and GPTBot (-0.5 pp), with a smaller gain to the search tier (Claude-SearchBot +0.6 pp). Where May's gains concentrated in Bytespider, June's concentrated — even harder — in ClaudeBot.
| AI Bot | May 2026 Share | June 2026 Share | Change (pp) | Relative Change |
|---|---|---|---|---|
| Googlebot | 26.7% | 24.9% | -1.8 | -6.7% |
| ClaudeBot | 12.1% | 20.0% | +8.0 | +66.3% |
| Meta-ExternalAgent | 12.6% | 10.2% | -2.4 | -19.3% |
| GPTBot | 10.1% | 9.6% | -0.5 | -5.0% |
| Bingbot | 8.2% | 8.1% | -0.1 | -1.2% |
| Bytespider | 10.1% | 7.3% | -2.8 | -27.4% |
| Applebot | 6.9% | 5.8% | -1.0 | -15.1% |
| Amazonbot | 5.4% | 5.5% | +0.2 | +2.8% |
| Claude-SearchBot | 2.7% | 3.3% | +0.6 | +21.9% |
Six trends stand out from this comparison:
-
ClaudeBot is the story of the month — the biggest surge on record. Anthropic's training crawler jumped from 12.1% to 20.0% (+8.0 pp, +66% relative) — the single largest absolute and relative gain any crawler has posted since I began tracking. ClaudeBot passed Meta-ExternalAgent and GPTBot in one month to become the #2 AI crawler on the web, and it now out-crawls the next two crawlers (Meta at 10.2% and GPTBot at 9.6%) combined. Anthropic is crawling at a scale that, a quarter ago, only Google operated at.
-
Bytespider's surge reversed. Last month's biggest gainer became this month's biggest structural loser. Bytespider fell -2.8 pp to 7.3%, dropping from #4 to #6. May's spike now looks like the same kind of one-to-two-month training burst that Applebot showed in April — a steep ramp that didn't hold. This is the second consecutive month a "hot crawler" cooled sharply, and the clearest recurring lesson in this report: momentum spikes revert.
-
ClaudeBot overtook everyone but Google. ClaudeBot's climb is the mirror image of last month, when it slipped below GPTBot. Anthropic's training crawler didn't just recover — it more than doubled the field behind Googlebot. And unlike Meta's finite Q1 campaign, ClaudeBot has been climbing on a V-shaped trajectory all year, which makes this surge harder to write off as a one-month burst (though the reversal lesson above urges caution).
-
Googlebot fell below 25%. After a -1.8 pp drop to 24.9%, Googlebot broke the 25% floor for the first time. Google has now lost roughly 14 percentage points of AI-crawler share since January's ~39% — the most durable single trend in this dataset, and the one prediction that has hit every month.
-
The search tier keeps building, still led by Anthropic. Claude-SearchBot rose +0.6 pp to 3.3%, extending its lead as the largest dedicated AI search crawler. Anthropic now leads both the training tier (ClaudeBot) and the search tier (Claude-SearchBot) — a mirror image of OpenAI, whose GPTBot leads training but whose OAI-SearchBot trails in the tail. If you're building apps that rely on real-time retrieval, understanding what a web search API is and how these bots work under the hood matters more than ever.
-
Meta kept contracting. Meta-ExternalAgent fell a third straight month, -2.4 pp to 10.2%. The "plateau at 18-20%" I predicted in the Q1 review is now decisively wrong in the opposite direction — Meta has lost 6.5 pp across the quarter, confirming its aggressive Q1 Llama-training crawl was a finite campaign, not a new baseline.
What Are AI Bots Actually Doing With the Content They Crawl?
June's crawl-purpose data has one clear driver: ClaudeBot. Because ClaudeBot is a pure training crawler, its +8 pp surge pushed Training up +2.4 pp to 47.4% — resuming the climb after last month's plateau — while Mixed Purpose fell -2.9 pp to 38.8%, tracking Googlebot's decline as it has all year. And Search crossed 10% for the first time, rising +0.5 pp to a record 10.5%.
| Crawl Purpose | June 2026 Share | May 2026 Share | Change (pp) |
|---|---|---|---|
| Training | 47.4% | 45.0% | +2.4 |
| Mixed Purpose | 38.8% | 41.8% | -2.9 |
| Search | 10.5% | 10.0% | +0.5 |
| User Action | 2.5% | 2.7% | -0.1 |
| Undeclared | 0.8% | 0.6% | +0.1 |
A note on the Training figure: last month's edition reported Training at 51.8%. Cloudflare Radar re-based its crawl-purpose classification corpus in mid-Q2, so this edition computes both columns above from the current methodology (the
28dand28dControlwindows) rather than splicing to last edition's published number. The month-over-month delta (+2.4 pp) is clean; the absolute level shifted because the underlying classification changed, not because crawling collapsed. See the methodology note at the end.
Here's what the June data tells website owners:
Training resumed its climb to 47.4%. After stalling last month, Training gained +2.4 pp — and the cause is legible in the crawler table: ClaudeBot, which does nothing but collect training data, nearly doubled. Training is still the clear plurality — for every 100 AI bot requests, roughly 47 are explicitly dedicated to training — but the composition is shifting. A year ago this share was dominated by Google and Meta; now a single lab (Anthropic) is the marginal mover. The land-grab isn't over; it has a new lead bidder.
Mixed Purpose kept sliding to 38.8%. Crawlers that simultaneously index for search and collect training data have now lost ground for six consecutive months, a decline that tracks Googlebot's own almost exactly. You still can't separate the AI training from the search indexing with these crawlers — block Googlebot and you disappear from Google Search. For developers looking at how to ground AI responses with Google Search alternatives, this bundling problem keeps coming up.
Search crawling passed 10% — a threshold I flagged in May. AI search crawlers rose to a record 10.5%, led again by Claude-SearchBot (now 3.3%, the largest dedicated AI search crawler). In the May edition I forecast Search would clear 10% in June; it did. This is the strongest signal yet that the AI industry's center of gravity is shifting from "crawl everything once to train" toward "fetch pages in real time to answer queries" — and it's the trend most relevant to whether your content gets cited rather than just ingested. The growing ecosystem of AI search API alternatives is the developer-facing side of exactly this shift.
User Action held near 2.5%. Real-time, user-triggered fetches (ChatGPT-User and equivalents) stayed roughly flat. This is the category most directly tied to AI assistants' "browse" features; it dipped a hair this month but remains structurally higher than at the start of the year.
The structural story of June is that retrieval kept its foothold even as training re-accelerated. Search set a record while Training climbed — the two aren't zero-sum, and the crawler that drove Training up (ClaudeBot) belongs to the same operator driving Search up (Claude-SearchBot). Anthropic is scaling both pipelines at once, which is exactly why it now sits second in total crawler share.
Which Industries Are AI Bots Targeting Most?
The industry mix moved again in June — and reversed last month's move. Shopping & General Merchandise rebounded from 25.3% to 27.2% (+2.0 pp), recovering most of the ground it lost in May, while Internet & Telecom (-1.2 pp) and Computer & Electronics (-0.7 pp) gave some back. The retail-content pendulum swung again, which is itself the pattern: no single vertical has held a durable trend this quarter.
| Industry Vertical | June 2026 Share | May 2026 Share | Change (pp) |
|---|---|---|---|
| Shopping & General Merchandise | 27.2% | 25.3% | +2.0 |
| Internet and Telecom | 20.3% | 21.5% | -1.2 |
| Computer and Electronics | 18.4% | 19.1% | -0.7 |
| News, Media, and Publications | 9.6% | 9.3% | +0.3 |
| Gambling | 6.6% | 6.6% | -0.1 |
| Business and Industry | 3.5% | 3.5% | -0.1 |
| Professional Services | 2.6% | 2.7% | -0.1 |
| Finance | 2.5% | 2.6% | -0.1 |
| Games | 2.2% | 2.3% | -0.1 |
Shopping is comfortably the most-crawled vertical again, restoring a near-7 pp lead over Internet & Telecom. The steadiest riser remains News, Media & Publications, up a third straight month to 9.6% — as Search and User-Action crawling grows (see the crawl-purpose section above), AI assistants are fetching more news and reference content in real time to answer time-sensitive queries, a different pattern than bulk e-commerce training scrapes. That slow, consistent climb is more informative than the month-to-month retail swings, and it's exactly the content type most exposed to the retrieval shift.
The industry-level breakdown gets more granular:
| Industry | June 2026 Share | May 2026 Share | Change (pp) |
|---|---|---|---|
| Retail | 24.7% | 22.8% | +1.9 |
| Computer Software | 16.7% | 17.3% | -0.6 |
| IT and Services | 6.2% | 7.0% | -0.8 |
| Gambling & Casinos | 6.1% | 6.1% | 0.0 |
| Marketing and Advertising | 5.2% | 5.3% | -0.2 |
| Adult Entertainment | 5.0% | 4.7% | +0.3 |
| Media | 4.8% | 4.9% | -0.1 |
| Internet | 4.7% | 4.7% | 0.0 |
| Telecommunications | 2.8% | 2.9% | -0.1 |
The most notable industry-level move is Retail rebounding +1.9 pp to 24.7%, mirroring the Shopping vertical's recovery, while Computer Software eased -0.6 pp to 16.7% but held its #2 spot — AI crawlers are still targeting documentation, code, and developer content heavily. Adult Entertainment ticked up +0.3 pp to 5.0%, edging ahead of Media, a recurring pattern when training crawlers do broad recrawls. If you maintain developer documentation, API references, or technical tutorials, this is why choosing the right AI web search API for your applications matters — these crawlers are the infrastructure behind the search results your users see, and software content remains the second-most-crawled industry on the web.
AI training crawlers aren't the only bots scanning the web at this scale, either. Technology detection platforms like Technologychecker.io crawl and fingerprint over 50 million domains using HTTP header analysis, JavaScript fingerprinting, DNS lookups, and headless browser rendering to identify 40,000+ technologies. Unlike AI training crawlers that take content for model weights, technology intelligence crawlers need to re-crawl frequently to track stack changes and new tech adoptions.
How Much Traffic Do AI Crawlers Give Back?
Cloudflare Radar's crawl-to-refer ratio — the number of pages an operator crawls for every one referral it sends back to a website — is the single most honest measure of whether a given AI company is a fair exchange or a pure extractor. June's numbers held their overall shape but produced one striking move: even as ClaudeBot's crawling nearly doubled, Anthropic's ratio improved more than 3× — from ~9,478:1 to ~2,978:1 — because its search crawler referred far more traffic than the extra training crawls cost it. It remains, by a wide margin, the most extractive operator on the list.
| Operator | June 2026 (crawls : 1 referral) | May 2026 | Direction |
|---|---|---|---|
| Anthropic | 2,978 : 1 | 9,478 : 1 | Most extractive (improving) |
| OpenAI | 530 : 1 | 945 : 1 | Improving |
| Mistral | 302 : 1 | 69 : 1 | Worsening fast |
| Perplexity | 215 : 1 | 163 : 1 | Worsening |
| Microsoft | 37 : 1 | 33 : 1 | Slightly worse |
| Yandex | 25 : 1 | 25 : 1 | Flat |
| Baidu | 11 : 1 | 12 : 1 | Improving |
| ByteDance | 11 : 1 | 9 : 1 | Slightly worse |
| 5 : 1 | 5 : 1 | Most generous | |
| DuckDuckGo | 2.3 : 1 | 1.7 : 1 | Near-even |
According to Cloudflare Radar, Anthropic still crawled roughly 2,978 pages in June for every single visitor it referred back — an improvement from last month's ~9,478:1, but still an order of magnitude more lopsided than any other major operator, consistent with ClaudeBot's training-heavy footprint. OpenAI improved to 530:1, while Google returns the most traffic relative to what it takes, at just 5:1 — the structural advantage of running search and AI crawling through bundled infrastructure that still drives clicks. DuckDuckGo, which leans on others' indexes, is nearly even at 2.3:1.
For website owners, this is the number that reframes the "should I block AI crawlers?" question. A 5:1 ratio (Google) is a recognizable search-engine bargain — you give crawl access, you get visitors. A ~3,000:1 ratio (Anthropic) is not a bargain in the traditional sense; it's content acquisition with very little traffic return. That gap is the clearest data-backed case for treating training crawlers and search crawlers differently in your robots.txt — exactly the separation Anthropic's own ClaudeBot/Claude-SearchBot split now makes possible. And the direction of travel matters: Anthropic's ratio is improving precisely because its search crawler (which cites and refers) is growing alongside its training crawler (which doesn't) — a preview of the E-E-A-T argument below.
Source: Cloudflare Radar — radar/bots/crawlers/summary/crawl_refer_ratio (radar.cloudflare.com), June 8 – July 6, 2026 vs. May 11 – June 8, 2026.
How Are Websites Fighting Back Against AI Crawlers?
⚠️ Methodology note: This section reports raw domain counts — the number of distinct domains whose robots.txt explicitly names each crawler — rather than percentages, to sidestep the share-denominator problem. One caveat carries extra weight this month: nearly every crawler's count rose +40 to +90 domains at once, which almost certainly reflects growth in Cloudflare's parsed robots.txt sample rather than a synchronized blocking wave. So read rank and divergence, not the absolute deltas — the meaningful signals are which crawlers moved against the rising tide and which climbed faster than it. June figures are the most recent Radar snapshot; May figures are last month's snapshot.
Most Referenced AI Crawlers in robots.txt — June 2026
| AI Crawler | Domains (June 2026) | Domains (May 2026) | Change | Operator |
|---|---|---|---|---|
| GPTBot | 712 | 632 | +80 | OpenAI |
| ClaudeBot | 623 | 539 | +84 | Anthropic |
| Google-Extended | 593 | 502 | +91 | |
| CCBot | 575 | 501 | +74 | Common Crawl |
| Bytespider | 493 | 420 | +73 | ByteDance |
| meta-externalagent | 441 | 371 | +70 | Meta |
| PerplexityBot | 422 | 382 | +40 | Perplexity |
| Amazonbot | 419 | 367 | +52 | Amazon |
| Applebot-Extended | 409 | 338 | +71 | Apple |
| Googlebot | 393 | 402 | -9 | |
| ChatGPT-User | 369 | 346 | +23 | OpenAI |
| facebookexternalhit | 350 | 355 | -5 | Meta |
| OAI-SearchBot | 304 | 282 | +22 | OpenAI |
| anthropic-ai | 255 | 246 | +9 | Anthropic |
Against a tide where almost everything rose, two agents fell — and both are search-indexing bots. Googlebot (-9) and facebookexternalhit (-5) are the only entries to lose references while the parsed corpus grew, a clean signal that site owners are removing generic search-crawler disallows even as they add AI-specific ones. That's the whole game in two data points: keep the crawlers that send visitors, cut the ones that only take.
Three climbers beat the rising tide. Google-Extended (+91) widened its lead for #3, the fastest riser on the board — owners are opting out of Google's AI training (Google-Extended) while keeping search visibility (Googlebot, which fell). ClaudeBot (+84) posted the second-largest gain, tracking its traffic surge almost exactly — blocking follows visibility in logs. And Applebot-Extended jumped +71 to #9 (from #12), despite Applebot's traffic falling this month: robots.txt adoption lags traffic, so owners who noticed Apple during its April–May run are still adding directives now that the surge has faded. Adoption is sticky; it accumulates on awareness and rarely reverses.
The meta-externalagent gap persists: despite being the #3 AI crawler by traffic at 10.2%, Meta's bot ranks only #6 by robots.txt references — still one of the larger disparities between traffic share and blocking attention. The bigger new gap is ClaudeBot: now the #2 crawler by traffic at 20.0%, yet only #2 by references with a much smaller lead than its traffic would imply. If you maintain a robots.txt and you're worried about training crawlers, ClaudeBot and Meta-ExternalAgent are the two highest-traffic bots most under-referenced relative to their footprint this month.
Two qualitative observations continue to hold:
- Selective adoption, not blanket blocks. Website owners add rules for specific crawlers as they notice them, not site-wide AI bans. Google-Extended's rise and Applebot-Extended's stickiness both reflect awareness-driven, crawler-by-crawler decisions.
- The blocking cohort is still small. The top-referenced crawler appears in 632 domains' robots.txt in Radar's parsed sample — meaningful and growing steadily (+35 month-over-month), but a tiny fraction of the global web. The overwhelming majority of sites still allow every AI crawler unrestricted access.
What's Happening on Cloudflare Workers AI?
Beyond crawlers, Cloudflare Radar tracks usage patterns on Cloudflare Workers AI — the platform that lets developers run AI models at the edge. This month the metric didn't just reshuffle, it changed character: measured as share of inference requests, Workers AI is now overwhelmingly an embeddings platform. A single model — BAAI's multilingual BGE-M3 embedding model — accounts for 74.8% of all inferences, and Text Embeddings as a task category surged to 82.4%. The chat-model leaderboard that dominated earlier editions (Llama vs Kimi vs Gemma) now lives almost entirely in the long tail.
⚠️ Read this metric with care. Unlike the crawler shares above, Workers AI figures are share of inference requests, so a handful of high-volume batch jobs (bulk embedding a corpus, transcribing an audio archive) can dominate the mix and swing it hard month to month. Last month's control window was speech-dominated — Whisper alone was 47.6% of inferences and Automatic Speech Recognition was 49.3% of tasks — and that batch has since ended. Treat the table below as a snapshot of what ran, not a durable popularity ranking of models.
Most-Used Models (by Share of Inferences)
| Model | June 2026 Share | May 2026 Share | Task | Developer |
|---|---|---|---|---|
| BGE-M3 | 74.8% | 28.7% | Text Embeddings | BAAI |
| Gemma 4 26B-A4B-IT | 4.5% | 3.2% | Text Generation | |
| BGE Base EN v1.5 | 3.8% | 4.0% | Text Embeddings | BAAI |
| Whisper Large V3 Turbo | 2.1% | 47.6% | Speech (ASR) | OpenAI |
| BGE Small EN v1.5 | 2.1% | 3.3% | Text Embeddings | BAAI |
| Qwen3 30B A3B | 1.4% | 1.3% | Text Generation | Alibaba |
| M2M-100 1.2B | 1.3% | ~1% (tail) | Translation | Meta |
| Llama 4 Scout 17B | 1.3% | 1.0% | Text Generation | Meta |
| EmbeddingGemma 300M | 0.9% | ~1% (tail) | Text Embeddings |
The headline is retrieval, not generation. Four of the nine top models are embedding models (BGE-M3, BGE Base, BGE Small, EmbeddingGemma), and together they dwarf everything else — BGE-M3 alone went from 28.7% to 74.8% of inferences as one or more large vector-indexing workloads scaled up on the edge. Embeddings are the engine of vector search and RAG, so this is the same "shift toward retrieval" the crawler data has been telling all year, showing up on the inference side: developers are running the edge to build search indexes, not just to chat.
The chat-model race moved to the tail — and produced another reversal. Among text-generation models, Google's Gemma 4 26B is now the leader at 4.5%, ahead of Alibaba's Qwen3 30B (1.4%) and Meta's Llama 4 Scout (1.3%). Notably, Moonshot AI's Kimi K2.6 — last month's breakout at ~0.9% and climbing — fell out of the top nine entirely. Just as Applebot's and Bytespider's traffic spikes didn't hold, Kimi K2.6's momentum didn't either. It's the third "hot-then-reverted" data point this edition, on the third different dataset — the clearest cross-cutting lesson of the month.
Task Distribution
| Task Type | June 2026 Share | May 2026 Share | Change (pp) |
|---|---|---|---|
| Text Embeddings | 82.4% | 38.1% | +44.3 |
| Text Generation | 11.4% | 9.2% | +2.3 |
| Automatic Speech Recognition | 2.8% | 49.3% | -46.5 |
| Translation | 1.3% | 0.8% | +0.5 |
| Text Classification | 1.0% | 1.3% | -0.4 |
| Text-to-Image | 0.9% | 1.2% | -0.3 |
Text Embeddings surged +44.3 pp to 82.4% while Automatic Speech Recognition collapsed -46.5 pp to 2.8% — the two moves are effectively a single story: May's speech-transcription batch ended and a much larger embedding workload took its place. Text Generation held steady around 11%. Look past the batch-driven swings and the durable signal is consistent with the crawler data — the edge is increasingly used to power retrieval (embeddings feeding vector search and RAG) rather than one-off generation. For a primer on why embeddings sit under modern AI search, our explainer on how to build AI search for a website walks through the retrieval pipeline these workloads feed.
What the June Data Means for E-E-A-T and AI Citability
Two numbers in this month's data change what "optimizing for AI crawlers" should mean. Search crawling hit a record 10.5%, and Anthropic's Claude-SearchBot — the crawler that fetches a page to answer a live question and cite a source — is now the largest dedicated AI search crawler at 3.3%. The crawl-to-refer data makes the stakes concrete: search crawlers are the ones that send visitors back, while pure training crawlers (Anthropic's ClaudeBot still crawls ~3,000 pages per referral) take content and return almost nothing. The game is shifting from being scraped to being cited — and getting cited is precisely what Google's E-E-A-T framework (Experience, Expertise, Authoritativeness, Trustworthiness) was built to earn. The same signals that make a human editor trust a page are what make a retrieval crawler quote it.
Here's how each June signal maps to an E-E-A-T lever you can actually pull:
| E-E-A-T pillar | June signal driving it | What it rewards |
|---|---|---|
| Experience | Search crawling at a record 10.5%; retrieval fetches specifics | First-hand data, screenshots, tests, lived detail an AI can quote |
| Expertise | Training up to 47.4% as ClaudeBot re-crawls for refresh | Depth, accurate terminology, named authors with real credentials |
| Authoritativeness | News, Media & Publications up a third straight month (9.6%) | Being the primary source others cite; resolvable identity |
| Trustworthiness | Crawl-to-refer gap — search crawlers reward citable pages | Verifiable claims, dated data, disclosed methodology and authorship |
Experience — the retrieval shift rewards first-hand specifics. Claude-SearchBot and its peers fetch pages to answer a specific question, then decide what to quote. Generic, paraphrased content loses that competition to pages carrying something only the author could produce: original measurements, screenshots, reproducible tests, real numbers. This very report is the pattern in miniature — it's citable because it runs original queries against the Cloudflare Radar API each month, not because it restates someone else's summary. If you want AI answers to pull from you, publish the thing that can't be paraphrased from elsewhere.
Expertise — training crawlers re-crawl for refresh, so depth compounds. Training climbed +2.4 pp this month because ClaudeBot re-crawled at scale, and re-crawling is a refresh behavior: models re-ingest content that is updated, technically precise, and attributable to someone who demonstrably knows the field. Thin, undated, anonymous content decays out of the refresh cycle; deep content with a named, credentialed author (see this report's author bio and credentials block) keeps getting picked back up. Expertise isn't a meta tag — it's demonstrated in the substance and the byline.
Authoritativeness — the one steady riser is News, Media & Publications. Amid retail's month-to-month swings, publication content has now grown three months straight, exactly as real-time retrieval scales. AI assistants answering time-sensitive questions reach for authoritative, well-sourced pages — so the lever is to become the source other pages cite, and to make your identity machine-resolvable. Clear Organization, Author, and Article structured data lets a crawler connect a claim to a credible entity; our primer on how search engines really work covers why entity and link signals still anchor authority in an AI-retrieval world.
Trustworthiness — the crawl-to-refer gap is a trust ledger. The reason search crawlers refer traffic and training crawlers don't is that a search crawler will only surface a page it can quote confidently. Make that easy: cite primary sources, date every statistic, show your methodology (this report names its exact API endpoints and 28-day windows), and disclose who wrote it. Those are the signals that let a retrieval system attach your claim to a citation instead of dropping it. For developers building on the other side of this — grounding AI answers in verifiable sources — what a web search API is walks through the retrieval layer that makes citation possible.
💡 Expert Insight: The reversal lesson from the crawler data applies to E-E-A-T too. The bots that spiked and faded this quarter (Applebot, Bytespider, Kimi K2.6) are a warning against spiky tactics. Durable AI citability is built the same way durable authority is — first-hand experience, demonstrated expertise, earned references, and verifiable trust signals accumulated over months. The retrieval shift means those signals now pay off in AI answers, not just blue links.
The through-line: as crawling tilts from training toward retrieval, the content that wins is the content an AI can stand behind when it cites you. Every action in the website-owner checklist below — separating search from training crawlers, allowing Claude-SearchBot, keeping content fresh — is downstream of one idea: make your pages the ones AI systems trust enough to quote.
Q1 2026 Quarterly Review: The Great Reshuffling
Q1 2026 was the most transformative quarter in the history of AI web crawling. This section compiles data from three consecutive monthly analyses to document the structural shifts reshaping how AI companies interact with the open web.
Q1 2026 at a Glance
| Metric | January 2026 | February 2026 | March 2026 | Q1 Change |
|---|---|---|---|---|
| Top crawler (Googlebot) | 38.7% | 34.6% | 31.6% | -7.1 pp |
| #2 crawler | GPTBot (12.8%) | Meta-ExternalAgent (15.6%) | Meta-ExternalAgent (16.7%) | Meta took #2 |
| Training crawl share | 42.0% | 45.4% | 49.9% | +7.9 pp |
| Mixed Purpose crawl share | 48.3% | 43.9% | 39.9% | -8.4 pp |
| Top 5 company concentration | 84.5% | 82.8% | 80.2% | -4.3 pp |
| Domains blocking GPTBot | 5.29% | 5.45% | 5.52% | +0.23 pp |
| Llama 3 8B (Workers AI) | 41.7% | 40.1% | 37.3% | -4.4 pp |
How Each Crawler Moved Across the Quarter
| AI Bot | Jan 2026 | Feb 2026 | Mar 2026 | Q1 Change (pp) | Q1 Direction |
|---|---|---|---|---|---|
| Googlebot | 38.7% | 34.6% | 31.6% | -7.1 | Declining |
| Meta-ExternalAgent | 11.6% | 15.6% | 16.7% | +5.1 | Rising |
| GPTBot | 12.8% | 12.1% | 12.0% | -0.8 | Stable/Slow decline |
| ClaudeBot | 11.4% | 11.1% | 11.7% | +0.3 | V-shaped recovery |
| Bingbot | 9.7% | 9.3% | 8.2% | -1.5 | Declining |
| Applebot | 2.5% | 3.1% | 5.8% | +3.3 | Surging |
| Amazonbot | 4.8% | 5.4% | 4.4% | -0.4 | Volatile |
| Bytespider | 3.5% | 3.3% | 3.6% | +0.1 | Flat |
| OAI-SearchBot | 2.0% | 2.6% | 2.2% | +0.2 | Volatile |
Googlebot's sustained decline is the defining trend of Q1. Losing 7.1 pp over three months represents the largest quarterly share loss for any single crawler in Cloudflare Radar's tracking history. This doesn't necessarily mean Google is crawling less — it means competitors are crawling more, faster. At 31.6%, Googlebot is no longer 2x the size of the next-largest crawler; the ratio to Meta-ExternalAgent is now 1.9x and shrinking.
Meta-ExternalAgent was the biggest winner of Q1. Gaining +5.1 pp over three months, Meta's crawler overtook GPTBot in February and never looked back. The growth rate decelerated across the quarter (+3.1 pp → +3.7 pp → +2.3 pp), suggesting Meta may be approaching a near-term equilibrium.
Applebot's March surge was the quarter's biggest surprise. After modest growth in January (+0.2 pp) and February (+0.6 pp), Applebot exploded in March (+3.2 pp, +124% relative). Over Q1, Applebot more than doubled from 2.5% to 5.8%. Apple is now a top-six AI crawler operator — a status no one predicted at the start of the quarter.
The Training-Mixed Purpose Crossover
The most consequential structural shift of Q1 2026 happened in the crawl purpose data:
| Crawl Purpose | Jan 2026 | Feb 2026 | Mar 2026 | Q1 Change (pp) |
|---|---|---|---|---|
| Training | 42.0% | 45.4% | 49.9% | +7.9 |
| Mixed Purpose | 48.3% | 43.9% | 39.9% | -8.4 |
| Search | 6.9% | 8.2% | 7.7% | +0.8 |
| User Action | 2.2% | 2.0% | 2.1% | -0.1 |
| Undeclared | 0.5% | 0.4% | 0.4% | -0.1 |
In January, Mixed Purpose led Training by 6.3 pp (48.3% vs 42.0%). By February, Training had overtaken Mixed Purpose for the first time (45.4% vs 43.9%). By March, the gap had widened to 10 pp (49.9% vs 39.9%).
This crossover represents a fundamental change in how AI companies interact with the web:
- Before Q1 2026: The majority of AI bot traffic came from dual-purpose crawlers (Googlebot, Bingbot) that bundled search indexing with AI training. Website owners couldn't separate the two.
- After Q1 2026: The majority comes from purpose-built training crawlers (GPTBot, ClaudeBot, Meta-ExternalAgent, Applebot) that exist solely to collect AI training data. Website owners can block these selectively.
The shift gives website owners more control. You can now block the majority of AI training crawling without sacrificing search visibility — something that wasn't possible when Mixed Purpose crawlers dominated.
Market Concentration Is Declining
| Metric | January | February | March |
|---|---|---|---|
| Top 5 companies (Google, Meta, OpenAI, Anthropic, Microsoft) | 84.5% | 82.8% | 80.2% |
| Top 6 companies (+ Apple) | 87.0% | 85.9% | 86.0% |
| Top 4 crawlers share | 74.4% | 73.5% | 72.0% |
The AI crawler market is diversifying. The traditional top five lost 4.3 pp of concentration over Q1, but when you include Apple as a sixth major player, the top-six concentration remained remarkably stable at ~86%. The diversification is happening within the top tier, not from outside it.
Quarterly Robots.txt Blocking Trends
| AI Crawler | Jan 2026 | Feb 2026 | Mar 2026 | Q1 Change |
|---|---|---|---|---|
| GPTBot | 5.29% | 5.45% | 5.52% | +0.23 pp |
| CCBot | 4.40% | 4.63% | 4.53% | +0.13 pp |
| ClaudeBot | 4.33% | 4.62% | 4.72% | +0.39 pp |
| Google-Extended | 4.04% | 4.36% | 4.37% | +0.33 pp |
| Googlebot | 4.13% | 3.77% | 3.80% | -0.33 pp |
| Bytespider | 3.69% | 3.73% | 3.67% | -0.02 pp |
| meta-externalagent | 3.24% | 3.26% | 3.34% | +0.10 pp |
| Amazonbot | 2.99% | 3.16% | 3.29% | +0.30 pp |
ClaudeBot saw the largest blocking increase across Q1 at +0.39 pp, overtaking CCBot to become the second-most referenced AI crawler in robots.txt files behind GPTBot. The blocking wave is real but slow — over the entire quarter, GPTBot grew from 5.29% to just 5.52%. More than 94% of domains still allow all AI crawlers unrestricted access. At the current rate, it would take years for blocking rates to reach even 10%.
The Traffic-Blocking Gap
| AI Crawler | March Traffic Share | March Blocking Rate | Gap (pp) |
|---|---|---|---|
| Googlebot | 31.6% | 3.80% | 27.8 |
| Meta-ExternalAgent | 16.7% | 3.34% | 13.4 |
| GPTBot | 12.0% | 5.52% | 6.5 |
| ClaudeBot | 11.7% | 4.72% | 7.0 |
| Amazonbot | 4.4% | 3.29% | 1.1 |
Meta-ExternalAgent's 13.4 pp gap is the most actionable — it's a dedicated training crawler with no search indexing benefit, yet it's blocked by fewer domains than GPTBot despite generating more traffic.
Workers AI Model Ecosystem Evolution
| Model | Jan 2026 | Feb 2026 | Mar 2026 | Q1 Change |
|---|---|---|---|---|
| Llama 3 8B Instruct | 41.7% | 40.1% | 37.3% | -4.4 pp |
| Stable Diffusion XL Base 1.0 | 13.4% | 13.4% | 12.3% | -1.1 pp |
| Whisper | 8.5% | 8.3% | 7.5% | -1.0 pp |
| Llama 4 Scout 17B | 7.7% | 6.7% | 7.0% | -0.7 pp |
| M2M-100 1.2B | 5.6% | 5.4% | 5.1% | -0.5 pp |
| Llama 3 8B Instruct (AWQ) | 4.7% | 4.9% | 4.4% | -0.3 pp |
| FLUX.1 Schnell | 2.4% | 2.5% | 3.0% | +0.6 pp |
| GPT-OSS 120B | -- | 1.6% | 2.1% | New (+2.1 pp) |
| Whisper Large V3 Turbo | 1.4% | 1.6% | 1.7% | +0.3 pp |
Every top model except FLUX.1 Schnell, GPT-OSS 120B, and Whisper Large V3 Turbo lost share across Q1. Meta's dominance is slowly eroding — from ~60% in January to 53.8% in March. GPT-OSS 120B is the breakout model of Q1, debuting in February and climbing to 2.1% in March. FLUX.1 Schnell is gaining on Stable Diffusion XL, emerging as a credible challenger in the text-to-image space.
Five Structural Shifts That Defined Q1 2026
-
The Training Takeover. Training crawlers went from minority (42.0%) to near-majority (49.9%) in a single quarter. The web's content is now primarily being consumed for AI model weights rather than search indexing. Website owners who want to opt out of AI training have clearer tools to do so.
-
The Google Erosion. Googlebot lost 7.1 pp across Q1. This is almost certainly driven by competitors growing faster rather than Google crawling less, but the proportional shift matters for every website owner's traffic analysis.
-
The Meta Ascendancy. Meta-ExternalAgent gained +5.1 pp, overtook GPTBot for #2 in February, and finished the quarter at 16.7%. Meta's aggressive Llama training pipeline is driving unprecedented data collection.
-
Apple's Arrival. Applebot went from 2.5% to 5.8% across Q1, with most growth concentrated in March. Apple is now a top-six AI crawler operator and should be included in any website owner's AI bot management strategy.
-
The Diversification of Workers AI. The model ecosystem became meaningfully more diverse. Llama 3 8B's share fell from 41.7% to 37.3% as developers adopted newer models and the long tail grew from ~15% to ~20%.
Q2 2026 Predictions (Made in the April Edition)
Based on Q1 trends, here's what I forecast for Q2 2026 — scored against end-of-Q2 data in the tracker below:
-
Training crawlers will exceed 55%. The +7.9 pp Q1 trajectory suggests Training could reach 55-57% by June 2026, with Mixed Purpose falling below 35%.
-
Googlebot will drop below 30%. The current quarterly decline rate would put Googlebot in the 28-30% range by end of Q2.
-
Applebot will enter the top five. If Apple sustains even half of March's growth rate, Applebot could pass Bingbot (currently 8.2%) by May or June.
-
Meta-ExternalAgent will plateau around 18-20%. The deceleration trend suggests Meta's growth rate is stabilizing.
-
GPT-OSS 120B will enter the Workers AI top five. At its current growth rate, it could reach 3-4% by June, approaching FLUX.1 Schnell territory.
-
Robots.txt blocking will remain below 6% for all crawlers. The slow growth rate means meaningful blocking thresholds remain years away.
Q2 Predictions Tracker: End-of-Q2 Scorecard
June closes Q2, so this is the final scorecard for the six predictions made back in April — plus a check on the four revised calls I made last month for June specifically. The through-line of the whole quarter is now unmistakable: structural trends compounded; momentum spikes reversed. Every prediction grounded in a multi-month trend hit; every prediction extrapolating a single steep month missed.
The six original April predictions, scored against June:
| # | Q2 Prediction (made in April) | Window | June 2026 Actual | Verdict |
|---|---|---|---|---|
| 1 | Training crawlers will exceed 55% | By June 2026 | 47.4% (corpus re-based; never neared 55%) | ❌ Miss |
| 2 | Googlebot will drop below 30% | End of Q2 | 24.9% — below 25%, still falling | ✅ Hit (overshot) |
| 3 | Applebot will enter the top five | May or June | Hit in April, reversed; #7 at 5.8% in June | ⚠️ Hit, then reversed |
| 4 | Meta-ExternalAgent will plateau at 18-20% | Q2 | Fell to 10.2% (third straight decline) | ❌ Miss |
| 5 | GPT-OSS 120B will enter Workers AI top five | By June | Spiked in April, gone from the top tier by June | ⚠️ Hit, then reversed |
| 6 | Robots.txt blocking will remain below 6% | Q2 | Top crawler in 712 of thousands of domains | ✅ Holds |
And the four revised calls I made last month for June:
| # | June call (made in May edition) | June 2026 Actual | Verdict |
|---|---|---|---|
| A | Bytespider will challenge GPTBot for #3 | Bytespider reversed to 7.3% (#6); didn't | ❌ Miss |
| B | Kimi K2.6 will challenge Llama 3 8B for #1 | Kimi fell out of the top 9; metric went to embeddings | ❌ Miss |
| C | Search crawl purpose will pass 10% | Reached a record 10.5% | ✅ Hit |
| D | Training will stay flat (not reach 55%) | Did not reach 55% (47.4%); climbed modestly | ✅ Hit |
The two calls that reversed both bet on momentum. I predicted Bytespider would keep climbing to challenge GPTBot (A) and Kimi K2.6 would keep climbing toward #1 (B). Both had posted steep one-month gains; both gave those gains back. That's now four momentum-based reversals across this quarter — Applebot, GPT-OSS 120B, Bytespider, Kimi K2.6 — spanning three different datasets. If there's one durable rule this report has earned, it's this: a single steep month is a spike until it survives a second month.
The structural calls all hit. Googlebot's sub-30% prediction (#2) overshot to 24.9% with no floor in sight. Search clearing 10% (C) landed on schedule. Training not reaching 55% (D) held. Each was grounded in a trend that had already run for months, not a one-month jump.
Forecasts for Q3 (July–September):
- ClaudeBot's surge will decelerate — watch for a partial reversal. By the reversal rule above, a +8 pp month is exactly the kind of spike that tends to retrace. My base case: ClaudeBot holds #2 but gives back some share, rather than continuing straight at Googlebot's #1.
- Googlebot will fall below 24%. The single most reliable trend in this dataset; nothing has arrested it.
- Search crawl purpose will hold above 10% and Claude-SearchBot will pass 4%. Retrieval is the one structural riser, not a spike — expect it to keep grinding up.
- Anthropic will remain the #1 operator by combined crawl share (ClaudeBot + Claude-SearchBot), even if ClaudeBot alone slips — because its search crawler keeps growing independently of the training surge.
What Should Website Owners Do About June's Trends?
Based on what I've found in the June 2026 Cloudflare Radar data, here are the actions I'd prioritize:
Make ClaudeBot your top robots.txt priority this month. Anthropic's ClaudeBot nearly doubled to 20.0% — the #2 AI crawler on the web, behind only Googlebot, and now the single fastest-growing major crawler. If your robots.txt was tuned around Google, Meta, and OpenAI, it now under-weights the bot generating one in five AI crawler requests. Decide deliberately how you want to treat it (see the training-vs-search point below) and add an explicit ClaudeBot directive rather than leaving it to a catch-all.
Don't over-react to last month's Bytespider surge. I told you in May to make Bytespider your top priority because it had surged to #4. It has since fallen back to #6 (10.1% → 7.3%) — the same reversal Applebot showed a month earlier. Keep any Bytespider directive you added; it costs nothing and ByteDance may ramp again. But the urgent case this month is ClaudeBot, not Bytespider. When a crawler spikes, wait for a second month before treating the spike as the new normal.
Separate training from search — the crawl-to-refer data proves why. Anthropic still crawls roughly 3,000 pages for every referral it sends back (vs. Google's 5:1), even after a big improvement. Claude-SearchBot (search, now 3.3% and the largest AI search crawler) is distinct from ClaudeBot (training, 20.0%). If you want Claude's search to surface and cite your content while opting out of bulk training, use separate directives — allow Claude-SearchBot, decide on ClaudeBot — and apply the same logic to OpenAI's GPTBot/OAI-SearchBot split.
Optimize for the retrieval shift, not just against training. Search crawling hit a record 10.5% and the largest search crawler (Claude-SearchBot) keeps growing. A robots.txt written purely to block bulk training scrapers will increasingly miss the real-time search and user-action fetches that actually drive AI-assistant citations back to your site. Decide deliberately whether you want to be in AI answers (allow search crawlers, and make your content easy to cite — see the E-E-A-T section) or out of training sets (block training crawlers). They're now genuinely separable.
Don't assume Meta-ExternalAgent is still a top threat. Meta's bot fell a third straight month to 10.2% (from 16.7% in March). Its aggressive Q1 Llama-training crawl was a finite campaign, not a permanent baseline — though at 10.2% it's still the #3 crawler and remains under-referenced in robots.txt relative to its traffic, so a directive is still worth having.
Audit your Googlebot assumptions. Googlebot fell below 25% to 24.9%, down from ~39% in January. For every 10 AI bot requests hitting your site at the start of the year, Google accounted for ~4; now it's closer to ~2.5. The rest are increasingly dedicated training and search crawlers — led this month by a single operator, Anthropic — not search-indexing bots.
How WebSearchAPI.ai Fits Into the AI Crawler Ecosystem
Every AI crawler in this report exists because AI companies need fresh, structured web data to power their models and search products. At WebSearchAPI.ai, we sit on the other side of this equation — providing developers and AI agents with a clean, fast, and affordable way to access real-time web data without running their own crawlers.
Here's why this matters in the context of May's trends:
- The balance keeps tilting toward real-time retrieval. Search crawling hit a record 10.5% and Claude-SearchBot extended its lead as the largest dedicated AI search crawler — and on Workers AI, embeddings (the engine of vector search and RAG) surged to 82.4% of inferences. Both signals point the same way: the next phase of AI is real-time fetching and retrieval, not just bulk training. Instead of building and maintaining crawling infrastructure for either job, WebSearchAPI.ai gives you instant access to structured search results, content extraction, and real-time web intelligence through a single API call.
- The crawler landscape reshuffles every month. With Anthropic's ClaudeBot surging to #2, Bytespider reversing, Anthropic leading the search tier, and the "hot crawler" turning over month after month, the complexity of managing web data access never settles. WebSearchAPI.ai handles the retrieval layer so you can focus on your application logic.
- Sub-second latency, 99.9% uptime, and structured responses mean your AI applications get the data they need without the infrastructure headaches that come with managing crawler fleets.
If the data in this report tells you anything, it's that the volume and complexity of AI web crawling is only accelerating — and the balance is now tilting toward real-time retrieval. WebSearchAPI.ai is purpose-built for developers who want to harness that web intelligence without becoming a crawling operation themselves. Learn more about what a web search API can do for your stack.
Frequently Asked Questions
What is an AI crawler?
An AI crawler (also called an AI bot or AI spider) is an automated program that visits websites to collect content for training artificial intelligence models or powering AI-powered search features. Unlike traditional search engine crawlers that index pages for search results, AI crawlers like GPTBot, ClaudeBot, and Meta-ExternalAgent specifically collect data to train large language models (LLMs). Some crawlers like Googlebot serve both purposes — indexing for search and collecting training data simultaneously. You can identify AI crawlers by their user-agent strings in your server logs or through tools like Cloudflare Radar.
How often is this AI crawler report updated?
This report is updated monthly with fresh data from Cloudflare Radar AI Insights. Each edition covers a rolling 28-day window and compares it against the immediately preceding 28-day window, so every month-over-month figure is computed on an identical basis. Quarterly editions add full-quarter trajectory analysis (the Q1 2026 review remains in this post as a historical anchor). Bookmark this page or check back at the beginning of each month for the latest analysis of AI crawler traffic patterns, market share shifts, and robots.txt directives.
Can I block AI crawlers from my website?
Yes. The primary method is adding disallow rules to your robots.txt file for specific AI crawler user agents. For example, adding User-agent: GPTBot followed by Disallow: / will request that OpenAI's crawler stop visiting your site. However, robots.txt is a voluntary protocol — crawlers are not technically required to obey it. As of June 2026, GPTBot and ClaudeBot remain the two most-referenced AI crawlers in robots.txt files (appearing in 712 and 623 domains respectively in Cloudflare Radar's parsed sample), with Google-Extended now third and rising fastest, but absolute adoption is still a small fraction of the web. Some CDN providers like Cloudflare also offer dashboard-level controls to block or rate-limit AI bots.
What is the difference between AI training crawlers and AI search crawlers?
AI training crawlers (like GPTBot, ClaudeBot, and Meta-ExternalAgent) collect web content to build and improve AI models. They typically scrape large volumes of content from many sites. AI search crawlers (like OAI-SearchBot and Claude-SearchBot) fetch specific pages in real time when a user performs a search query through an AI tool like ChatGPT or Claude. The key difference: training crawlers take your content to make the model smarter, while search crawlers fetch your content to answer a specific user question — and may drive traffic back to your site. As of June 2026, training crawling is the clear plurality at 47.4% of all AI bot traffic (up this month as Anthropic's training crawler ClaudeBot surged), while search crawling hit a record 10.5%, crossing the 10% line for the first time and led by Anthropic's Claude-SearchBot, now the largest dedicated AI search crawler at 3.3%.
Will blocking AI crawlers affect my SEO or search rankings?
Blocking dedicated AI training crawlers like GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider, or Applebot-Extended will not affect your rankings in Google, Bing, or other traditional search engines. These crawlers are separate from the search indexing bots. However, blocking Googlebot will remove your site from Google Search entirely since Google uses the same crawler for both search indexing and AI training. Google offers a middle ground with the Google-Extended user agent — blocking it opts you out of AI training while keeping your search presence intact, and it was the fastest-rising robots.txt directive in June 2026. Apple offers the same kind of separation with Applebot-Extended, which climbed into the top nine most-referenced robots.txt user agents as of June 2026.
Why did Bytespider's surge reverse in June 2026?
After Bytespider climbed from 3.6% (March) to 6.5% (April) to 10.1% (May) and briefly became the #4 AI crawler, it fell back to 7.3% in June, dropping to #6 behind Bingbot. The most likely explanation is that the spring ramp was a finite training burst rather than a sustained new baseline — the same pattern Apple's Applebot showed one month earlier. It's the recurring lesson of this report: a single steep month is a spike until it survives a second month. Website owners who added a Bytespider directive during the surge can keep it (it costs nothing and ByteDance may ramp again), but the urgent case in June is Anthropic's ClaudeBot, now the #2 crawler at 20.0%.
What changed with Anthropic's crawlers in June 2026?
Anthropic had the biggest month of any operator. ClaudeBot (training) surged from 12.1% to 20.0%, leapfrogging Meta-ExternalAgent and GPTBot to become the #2 AI crawler on the web — the largest single-month share gain any crawler has posted in this report's history. At the same time, Claude-SearchBot (search) rose to 3.3%, extending its lead as the single largest dedicated AI search crawler. Between the two, Anthropic now accounts for roughly 23% of all identified AI crawler traffic — second only to Google. This mirrors OpenAI's GPTBot/OAI-SearchBot split but in reverse: OpenAI leads in training, while Anthropic now leads in search (Claude's web search functionality) and is scaling training fastest. It gives website owners a clear reason to set separate directives for the two bots — allow the search crawler that may cite you, decide deliberately on the training crawler that won't.
How does Cloudflare track AI crawler traffic?
Cloudflare's global network spans 330+ cities in 125+ countries and processes over 81 million HTTP requests per second. Through its Radar platform, Cloudflare identifies and classifies AI bot traffic by analyzing user-agent strings, request patterns, and behavioral signatures across all sites on its network. The data in this report comes from Cloudflare Radar's AI Insights endpoints, which aggregate these signals into share-of-traffic percentages by bot, crawl purpose, industry, and region.
Which AI crawler is growing the fastest in 2026?
In June 2026, ClaudeBot had both the largest absolute and relative growth, +8.0 pp (+66%) (12.1% → 20.0%), making Anthropic's crawler the #2 AI bot on the web behind only Googlebot — the single largest one-month share gain any crawler has posted in this report's history. Bytespider, last month's standout, reversed sharply (-2.8 pp to 7.3%), and Claude-SearchBot was again the fastest riser in the search tier, climbing to 3.3% to extend its lead as the largest dedicated AI search crawler.
What percentage of web traffic comes from AI bots?
The percentages in this report represent share of identified AI bot requests, not share of total web traffic. Cloudflare Radar tracks the proportion of AI-related crawler activity relative to other AI bots, providing a competitive landscape view. The actual percentage of total web traffic from AI bots varies by website, but industry estimates suggest AI crawlers now account for a meaningful and growing share of overall internet traffic, particularly for content-heavy sites in retail, technology, and media.
How can I monitor AI crawler activity on my own website?
Check your server access logs for known AI bot user-agent strings (GPTBot, ClaudeBot, meta-externalagent, Applebot, Bytespider, Amazonbot, etc.). Most web analytics platforms filter out bot traffic by default, so log-level analysis gives the most accurate picture. Cloudflare users can view AI bot activity directly in their dashboard. For a structured approach, consider using a web search API to understand how your content appears in AI-powered search results and ensure your most important pages are properly accessible.
Cloudflare Radar Data Source & Methodology
Understanding where this data comes from — and what it can and cannot tell you — is critical for interpreting the trends above. Here's a full breakdown of how Cloudflare Radar collects, classifies, and aggregates the AI crawler data used in this report.
Network Scale
| Metric | Value |
|---|---|
| Global presence | 330 cities in 125+ countries |
| HTTP requests | 81 million/second average, peaks >129 million/second |
| DNS queries | 67 million/second (authoritative + resolver) |
This scale is what makes Cloudflare Radar one of the most comprehensive sources of internet traffic data available. The data in this report comes from two primary sources:
- Cloudflare's global network — real-time traffic data from HTTP requests flowing through their infrastructure
- 1.1.1.1 public DNS resolver — aggregated and anonymized DNS query data
For routing data, Cloudflare also uses RIPE RIS data from RIPE NCC (BGP route collectors).
How AI Bots Are Identified
Cloudflare uses a layered detection system to identify and classify AI crawlers:
- User-agent string matching — the most basic method; identifies bots that transparently announce themselves (GPTBot, ClaudeBot, etc.)
- Verified Bot Directory — manual approval process requiring bots to maintain public robots.txt commitments, use dedicated/verifiable IPs, unique user-agents, and honor crawl-delay settings
- Machine learning — supervised ML system that assigns a Bot Score (1-99)
- Heuristics — tailored rulesets for AI bot classification
- Behavioral analysis — pattern recognition from request sequences
- AI Labyrinth honeypot — hidden links to AI-generated decoy pages; bots that follow them are identified with high confidence since human visitors never see these links
- ai.robots.txt list — used as the basis for which AI bots to track
💡 Expert Insight: The layered approach matters because not all AI crawlers identify themselves honestly. User-agent matching catches transparent bots like GPTBot and ClaudeBot. Behavioral analysis and honeypots catch crawlers that disguise themselves as regular browsers.
How Crawl Purpose Is Classified
Bots are categorized into these purpose buckets:
- Training: dedicated training crawlers (GPTBot, ClaudeBot, Meta-ExternalAgent)
- Search: AI search bots (OAI-SearchBot)
- User Action: bots fetching pages for real-time user queries (ChatGPT-User)
- Mixed Purpose: bots serving dual roles like search indexing + AI training (Googlebot, Bingbot)
- Undeclared: purpose not identifiable
Data Aggregation Methods
- 7-day trailing average used to smooth daily fluctuations
- IPv4 addresses aggregated into /20 prefixes for visualization
- HTML traffic is separately classified into human, AI bot, and non-AI bot categories
- Normalization: data expressed as percentage of total requests (not absolute counts)
- Methodologies remain unchanged year-over-year for valid comparisons
- API data available under CC BY-NC 4.0 license
Caveats & Limitations
⚠️ Warning: Keep these limitations in mind when interpreting the data in this report:
- Countries with insufficient data volume are excluded from trend reporting
- Some metrics are available only at worldwide level, not per-country
- Mobile device categorization relies on User-Agent headers (accuracy limitations)
- Speed test data excludes locations with fewer than 100 tests per week
- The "location" filter corresponds to the billing country of the Cloudflare customer whose site received the traffic, not where the crawler is physically located
- Cloudflare sees traffic only to sites behind its network, not the entire internet, so the data is representative but not exhaustive
Report Parameters
This edition uses data from Cloudflare Radar's AI Insights endpoint (/radar/ai/bots/summary/*), Workers AI inference endpoint (/radar/ai/inference/summary/*), web-crawler endpoints (/radar/bots/crawlers/summary/{vertical|industry|crawl_refer_ratio}), and robots.txt analysis endpoint (/radar/robots_txt/top/user_agents/directive). The June 2026 monthly data covers the rolling 28-day window of June 8 through July 6, 2026, with every month-over-month comparison computed against the immediately preceding 28-day window (May 11 through June 8, 2026) via the API's dateRange=28d and dateRange=28dControl parameters, so both columns share an identical methodology. The Q1 2026 Quarterly Review compiles data from three consecutive monthly analyses covering January through March 2026.
I queried bot traffic breakdowns by user agent, crawl purpose, industry, and vertical; the new crawl-to-refer ratio by operator; Workers AI model and task distribution by account share; and domain-level robots.txt directives. All percentages represent share of identified AI bot requests (for crawling data) or share of accounts (for Workers AI data), not share of total web traffic. The crawl-to-refer ratio is a RATIO (crawls per referral), and robots.txt figures are raw domain counts.
Note on Revised Values
⚠️ Cloudflare Radar aggregates and may revise data after publication, and it re-based its crawl-purpose classification corpus in mid-Q2 2026. To keep month-over-month deltas clean, this edition computes both columns in every table from the current API windows (28d and 28dControl) rather than splicing to last edition's published figures, and reports robots.txt as raw domain counts rather than percentages. Two figures where last edition and this edition differ for methodological reasons: Training crawl share reads 47.4% here versus 51.8% last month (the classification corpus was re-based — the +2.4 pp month-over-month delta is the reliable number), and the prior-month ("May 2026") columns throughout are computed from the 28dControl window (May 11 – June 8), which is shifted ~6 days from last edition's published "May" window (May 5 – June 2). Where numbers differ slightly, the difference reflects Radar's revisions and window alignment, not a change in this report's method.
Data source: Cloudflare Radar AI Insights and Web Crawlers API endpoints (radar.cloudflare.com), June 8 – July 6, 2026 vs. May 11 – June 8, 2026. Last updated: July 6, 2026.
About the Author: I'm James Bennett, Lead Engineer at WebSearchAPI.ai, where I architect the core retrieval engine enabling LLMs and AI agents to access real-time, structured web data with over 99.9% uptime and sub-second query latency. With a background in distributed systems and search technologies, I've reduced AI hallucination rates by 45% through advanced ranking and content extraction pipelines for RAG systems. My expertise includes AI infrastructure, search technologies, large-scale data integration, and API architecture for real-time AI applications.
Credentials: B.Sc. Computer Science (University of Cambridge), M.Sc. Artificial Intelligence Systems (Imperial College London), Google Cloud Certified Professional Cloud Architect, AWS Certified Solutions Architect, Microsoft Azure AI Engineer, Certified Kubernetes Administrator, TensorFlow Developer Certificate.