Cloudflare-Perplexity tiff highlights how misguided AI agents are on the web
Perplexity says AI tools are different from traditional web searches, while Cloudflare points to obfuscation in forcing access to data even when websites say not to
An incipient battle between a content delivery network (CDN) and an artificial intelligence (AI) company may realign the contours of web standards, the idea of an open web, and how data is collected by AI companies.

Cloudflare, one of the world’s biggest CDNs fired the first shots, alleging that Perplexity AI is surreptitiously collecting data from websites that specifically instruct the AI firm’s bots to not do so.
A CDN is a distributed network of servers worldwide, for storing and delivering web content to users from a server closest to their geographic location, the advantages being of cost and performance.
The AI company countered by saying that the rise of AI-powered assistants and user-driven agents has blurred the boundary between what counts as “just a bot” and what serves the immediate needs of real people.
Perplexity said “companies like Cloudflare mischaracterise user-driven AI assistants as malicious bots”, but Cloudflare’s stand is that it is “giving content creators and publishers more control over how their content is accessed, which the company’s CEO Matthew Prince classified as AI’s existential threat to publishers.
To be sure, Cloudflare in June launched a service “pay per crawl” that allows companies that sign up to charge AI firms for crawling their websites. Prince said unless AI firms pay creators for their content, the default action would be to block AI crawlers.
Cloudflare noted that more than 2.5 million sites have opted to block AI training since July.
In this specific instance, Cloudflare said it has observed Perplexity crawling websites in ways that evade site owners’ preferences—stated explicitly in a website’s robots.txt file.
This file, hosted on a website’s front-end servers, is used to instruct web crawlers (such as search engines) which parts of the website they are allowed to access and list. To be sure, the instructions are directives that a crawler may or may not follow, not a piece of code that blocks a specific, or all, crawlers. The crawler bots, called user agents, usually declare their identity so that the site owners know who is crawling their content and can tailor their instructions accordingly.
The Robots Exclusion Protocol (of which robots.txt is a part) became an official standard under the Internet Engineering Task Force (IETF) with a proposed standard published in September 2022.
“We are observing stealth crawling behaviour from Perplexity, an AI-powered answer engine. Although Perplexity initially crawls from their declared user agent, when they are presented with a network block, they appear to obscure their crawling identity in an attempt to circumvent the website’s preferences,” Cloudflare detailed in a technical post.
“We received complaints from customers who had both disallowed Perplexity crawling activity in their robots.txt files and also created WAF rules (at the CDN level) to specifically block both of Perplexity’s declared crawlers: PerplexityBot and Perplexity-User. These customers told us that Perplexity was still able to access their content even when they saw its bots successfully blocked,” said Cloudflare.
Cloudflare’s tests, it said, were able to replicate this behaviour of obfuscation the company’s customer websites were complaining about. They also use a testing equivalence of contrast to make a further point — comparable tests with OpenAI’s ChatGPT crawler show that it stopped when disallowed, and did not follow up with other user agents after being blocked.
“I’ve been enjoying testing Comet (Peplexity’s browser) but this is not a good look for Perplexity. Cloudflare says Perplexity uses stealth crawling techniques, like undeclared user agents and rotating IP addresses, to evade robots.txt rules and network blocks,” said Glenn Gabe, of G-Squared Interactive, a Search Engine Optimization (SEO) consulting services firm.
Perplexity’s argument
Perplexity did not directly address the issue of potential obfuscation and bypassing instructions to access information from a website. The AI company said modern AI assistants are fundamentally different from traditional web crawling that was used by search engines over the years. It called the use of AI tools for search “user-driven” agents, which apparently don’t need to adhere to the conventional rules of the web.
“When you ask Perplexity a question that requires current information—say, ‘What are the latest reviews for that new restaurant?’—the AI doesn’t already have that information sitting in a database somewhere. Instead, it goes to the relevant websites, reads the content, and brings back a summary tailored to your specific question. This is fundamentally different from traditional web crawling, in which crawlers systematically visit millions of pages to build massive databases, whether anyone asked for that specific information or not,” the company said in a statement, in response to Cloudflare.
“Cloudflare says Perplexity is ignoring robots.txt and crawling sites using stock user agents. Anyone who’s been watching server logs already knew this was happening. Robots.txt was never a legal boundary or a real deterrent,” pointed out Brett Tabke, Founder of Pubcon, an SEO company.
Perplexity insisted that its user-driven agents do not store the information or train AI models with it. “When Google’s search engine crawls to build its index, that’s different from when it fetches a webpage because you asked for a preview. When Perplexity fetches a web page, it’s because you asked a specific question requiring current information,” Perplexity’s statement said.
The AI company said it appears Cloudflare confused Perplexity with 3-6 million daily requests of unrelated traffic from BrowserBase, “a third-party cloud browser service that Perplexity only occasionally uses for highly specialised tasks (less than 45,000 daily requests)”.
The way forward
Experts said AI chatbots are increasingly becoming default search tools, instead of traditional search engines such as Google Search and Microsoft Bing. Google too, within Search, is now increasingly layering AI, with add-ins such as AI Overviews, before listing relevant website links.
“I detect a ping pong match emerging between AI crawlers who have been ‘blocked’ as they’re scraping other people’s work, and those who are trying to protect the original work,” wrote Prof. Alan Woodward, of the Surrey Centre for Cyber Security, at the University of Surrey, in a post on X.
“How far will these AI bots go to harvest content?” he wondered.
The Cloudflare and Perplexity controversy represents a dilemma that publishers have been facing with the advent of commercial AI services. The traditional web allowed creators and publishers to be compensated through traffic monetisation, with search engine indexing sending users to their site and ad revenue or subscriptions to close the loop.
AI is breaking that chain, by directly serving answers, without a user being required to visit a website or click on a link.
As of last month, search giant Google faces an European Union antitrust complaint by a group of independent publishers, regarding AI Overviews in Search. Google’s AI Overviews are generated summaries related to any search, placed before traditional links to relevant websites or web pages.
“This is exactly why the CloudFlare bot blockers are so troubling,” added Tabke, citing specifics of the antitrust complaint — “Publishers using Google Search do not have the option to opt out from their material being ingested for Google’s Al large language model training and/or from being crawled for summaries, without losing their ability to appear in Google’s general search results page.”
Some experts said the bot behaviour is the norm rather than the exception.
“Cloudflare also has their own browser automation and scraping solution that anyone can use to scrape web pages that are disallowed by robots.txt,” pointed out Marcus Gill Greenwood, CEO of UBIO, an automation cloud platform.
“I actively want Perp to bypass robots.txt because they are doing so on my literal request. Totally different from model training and Cloudflare knows this,” writes Tyler Richards, data scientist at Streamlit, an open source framework for custom web apps.
ABOUT THE AUTHORVishal MathurVishal Mathur is Technology Editor for Hindustan Times. When not making sense of technology, he often searches for an elusive analog space in a digital world.

E-Paper


