TL;DR: What's changing between 2024 and 2026

AI crawlers are scraping at a historic rate while organic human traffic declines. Cloudflare documented an 18,000:1 crawl-to-1-click ratio at Anthropic in June 2025, and Similarweb measures a drop in Google-referred traffic to news sites. The CPM/subscription model is faltering.

Key points
  • Anthropic crawls 18,000 pages for 1 referred visitor (Cloudflare, June 2025).
  • AI Overviews appear on ~13.14% of Google queries (Semrush, 2025).
  • ChatGPT exceeds 800M weekly users in 2025 (OpenAI).
  • Organic CTR drops up to -34.5% when AI Overview is present (Ahrefs, 2025).

Why is AI bot traffic exploding in 2025-2026?

AI bot traffic quadrupled between Q1 and Q2 2025 on the TollBit network, increasing from one AI visit out of 200 to one out of 50, a 4x multiplication in six months (TollBit, State of the Bots Q2 2025). This acceleration follows the training of Claude, GPT-5, and Gemini.

Cloudflare observes the same trend on the infrastructure side. GPTBot, ClaudeBot, and PerplexityBot now represent a significant portion of non-human traffic on the 20% of the web protected by Cloudflare. Bytespider, TikTok's crawler, even surpasses Googlebot in volume on certain verticals.

Who are the main AI crawlers?

Four families dominate scraping in 2026. GPTBot and OAI-SearchBot for OpenAI, ClaudeBot and anthropic-ai for Anthropic, Google-Extended for Gemini, and PerplexityBot for Perplexity. ByteDance's Bytespider remains the most aggressive in raw volume. Each family has its own robots.txt rules and declared user-agent.

According to Cloudflare Radar, GPTBot tripled its query volume between late 2024 and mid-2025. Anthropic multiplied its volume by 5 over the same period, becoming the most active AI crawler on premium editorial sites.

Is human Google traffic really dropping due to AI Overviews?

Yes, and the data converges. An Ahrefs study of 300,000 keywords shows an average 34.5% drop in CTR for position 1 when an AI Overview is present (Ahrefs, 2025). Pew Research measures an even sharper drop for long informational queries.

Pew analyzed the behavior of 900 American users. When an AI Overview appears, only 8% click on a classic link, compared to 15% without an AI Overview, effectively halving outbound traffic (Pew Research, July 2025).

Which sites are most affected?

Similarweb published a ranking of the biggest declines in 2025. Definition and comparison sites like “what is X” are losing up to 40% of Google traffic year-on-year. Health sites Mayo Clinic and WebMD are declining. General how-to sites like WikiHow are seeing double-digit drops.

Conversely, strong brands and transactional sites are more resilient. Reddit exploded in organic visibility thanks to licensing deals with Google and OpenAI. Forums and UGC benefit from AI Overviews' preference for sources “with human opinion.” To optimize these new signals, see our on-page SEO expertise.

What is the actual scrape/click ratio of AI crawlers?

Cloudflare published the most telling figures in June 2025. Anthropic crawls 18,000 pages for every referred visitor. OpenAI shows a ratio of 1,500:1, Perplexity 70:1. Google remains at 6:1, the historical ratio for classic SEO (Cloudflare, 2025).

This imbalance is transforming the content economy. Previously, a publisher accepted that Googlebot consumed bandwidth because referred traffic largely compensated. With an Anthropic ratio of 18,000:1, server costs explode without advertising compensation. The promise of the open web is faltering.

Why is this gap so pronounced?

Three technical reasons explain it. LLMs respond directly without citing or linking in most cases. When they do cite, the click depends on the UX of each interface; ChatGPT does not always prominently display sources. Finally, autonomous agents scrape in loops to update their RAG indexes.

Profound, a GEO analytics platform, measures that only 6% of ChatGPT responses trigger an outbound click. On Perplexity, the rate rises to 24% thanks to the UX that pushes sources. This interface difference weighs more than the raw volume of queries.

What is the economic impact for publishers?

News Media Alliance estimates the value extracted annually by LLMs without direct remuneration to be several billion dollars (NMA, 2024). The Atlantic, BuzzFeed News, Vice have closed or pivoted in 2024-2025. The combined pressure of AI Overviews + scraping creates a brutal scissor effect on profitability.

On the revenue side, display CPM is declining on general information pages. Premium publishers are seeing their RPM drop by 10 to 25% depending on the vertical, mainly on queries now absorbed by AI Overviews. Subscriptions partially compensate, but only for very strong brands.

Are licensing deals saving the day?

Partially, and only for a select few. Reddit signed with Google for $60M/year and with OpenAI for an undisclosed amount. News Corp negotiated $250M over 5 years with OpenAI. The New York Times is suing OpenAI while discussing with other players.

For average publishers, access to these deals remains closed. Tollbit, ScalePost, and ProRata are developing pay-per-crawl marketplaces. Cloudflare launched its own pay-per-crawl program in 2025, allowing publishers to bill each scrape via an HTTP 402 response.

How to block or monetize AI crawlers?

Four options coexist in 2026. Purely block, conditionally allow, monetize, or ignore. Cloudflare activated default blocking of AI crawlers for new customers in July 2025, a major shift validated by 1 million domains according to their statement (Cloudflare, July 2025).

The robots.txt file remains the first line of defense. Disallow GPTBot, ClaudeBot, anthropic-ai, PerplexityBot, Google-Extended, Bytespider, CCBot. But 13.26% of AI queries ignore robots.txt according to TollBit. Technical blocking at the WAF or Cloudflare level becomes essential.

What is llms.txt and should it be adopted?

llms.txt is a proposal by Jeremy Howard published in September 2024. The file lists site content in a markdown format optimized for LLM ingestion. Adoption is heterogeneous in 2026, especially for technical docs (Anthropic, Mintlify, Stripe have implemented it). For publishers, its usefulness remains debated.

On the monetization side, the RSL (Really Simple Licensing) standard supported by News Media Alliance and Reddit proposes a machine-readable format for license terms. Coupled with Cloudflare's pay-per-crawl, it allows billing $0.001 to $0.01 per scrape. Our AI agency assists publishers with these decisions.

What new KPIs to track in 2026?

Clicks are no longer the sole metric. Profound, Peec.ai, and Otterly now measure the share of citations in ChatGPT, Claude, Perplexity, and Google AI Mode. The sectoral average for LLM share-of-voice fluctuates between 2 and 8% depending on the sector (Profound, 2025), equivalent to an SEO market share 10 years ago.

The KPIs to track are structured into three families. LLM visibility, citations in responses, share-of-voice per target prompt. Incoming AI traffic, sessions referred by chat.openai.com, perplexity.ai, claude.ai. Crawler costs, ratio of bytes served to bots versus human visitors, to be monitored in Cloudflare or server logs.

How to measure LLM citation share?

Three methods coexist. Periodic scraping of LLM responses via API on a basket of prompts, a method used by Profound and Peec. Tracking chat.openai.com and copilot.microsoft.com referrers in GA4, a passive and imperfect method. Server logs cross-referenced user-agent + IP, the most reliable but cumbersome method.

For classic SEO, optimization for AI Overviews remains possible. Structured content, FAQs, lists, and 40-60 word passages that answer an explicit question capture more citations. Our SEO agency tests these formats on dozens of sites.

FAQ: AI bot traffic and publisher strategy

What percentage of web traffic is non-human in 2026?

The Imperva Bad Bot Report 2024 estimates that 49.6% of web traffic is non-human, of which approximately 32% are bad bots. The share attributable to declared AI crawlers (GPTBot, ClaudeBot, etc.) remains minor in volume but is growing the fastest, with a quadrupling observed between Q1 and Q2 2025.

Should GPTBot and ClaudeBot be blocked now?

It depends on the business model. A site with high editorial value benefits from blocking or monetizing; Anthropic's 18,000:1 scrape/click ratio does not justify the bandwidth. A site seeking LLM visibility should allow GPTBot and OAI-SearchBot while monitoring costs.

Do AI Overviews affect all types of queries?

No. Semrush measures an average presence of 13.14%, but with huge variations by category. Health, legal, personal finance, and how-to queries exceed 30%. Transactional and navigational queries remain little affected. E-commerce suffers less from the AI Overview effect than news media.

How to appear in ChatGPT responses?

Three levers. Allow OAI-SearchBot in robots.txt to enter the live ChatGPT index. Produce 40-60 word content that directly answers a question, the format preferred by LLMs. Build strong topical authority; LLMs primarily cite sites with backlinks and extensive mentions.

How much does Cloudflare's pay-per-crawl generate?

Cloudflare does not publish average rates, as the program is in beta. Initial feedback suggests $0.001 to $0.01 per request depending on the content and publisher. For an average site receiving 100,000 monthly AI scrapes, the theoretical revenue is between $100 and $1000 per month, marginal but not zero.

What to remember for publishers in 2026?

The editorial ecosystem is shifting to a two-speed regime. On one side, strong brands sign 8-figure licensing deals and resist. On the other, average sites are hit by the double shock of AI Overviews + uncompensated scraping. Cloudflare and Similarweb figures leave no doubt about the direction of the movement.

The 2026 trade-off revolves around three axes. Deciding which crawlers to block, allow, or monetize. Measuring LLM citation share as seriously as SEO positions. Restructuring content for LLM extraction without degrading the human experience. These are topics we operate daily at Uclic.