For years, websites that did not want AI companies reading their content relied on a text file called robots.txt. It is a list of instructions that tells automated bots which pages they are allowed to visit. The problem is that it works entirely on trust. A bot can read the file and ignore it, and there is no enforcement mechanism. Patreon found out exactly how hollow that trust was: when it switched to active blocking, access attempts from individual AI training bots went from thousands per week to zero.
That gap, between asking and blocking, is the story here.
Patreon is using a tool from Cloudflare, a company that sits between most of the world's websites and the internet traffic trying to reach them. Cloudflare can identify bots by their behaviour patterns and cut off access before they ever reach the site's content. Crucially, Patreon is only blocking bots whose purpose is to train AI models. Bots that help users find content through search remain allowed, because those bots send real people back to creator pages.
The reason that distinction matters comes down to a simple exchange. For three decades, search engines crawled websites and in return sent human visitors back. It was a fair trade. AI training crawlers broke that trade. Cloudflare's data shows that Anthropic's crawler scrapes around 11,000 pages for every single visitor it sends back to a site. OpenAI's crawler is around 857 pages per referral. Google's search crawler sends back one visitor for roughly every five pages it reads. The difference is not marginal. It is structural.
Patreon is not alone in acting on this. Reddit struck licensing deals worth around $60 million per year with Google and a similar amount with OpenAI, then turned around and sued Anthropic for allegedly ignoring its instructions to stop scraping. The New York Times has an ongoing copyright lawsuit against OpenAI and Microsoft. Stack Overflow has a paid data deal with OpenAI. In September 2025, Reddit, Yahoo, and Medium backed a new open standard called Really Simple Licensing, which tries to create a common framework for how AI companies can license content from publishers, with options ranging from free access to pay-per-use arrangements.
Cloudflare itself has been building the infrastructure for this shift. In July 2025, it made AI bot blocking the default for new domains. It has since launched a pay-per-crawl system that lets websites charge AI companies for access, and a tool called AI Crawl Control that lets publishers set different rules for different types of bots. After Cloudflare switched on default blocking, its customers collectively blocked 416 billion AI bot scraping requests in the following five months.
What changed for Patreon specifically is that it recently introduced more publicly visible features, including a redesigned home feed and a short-form posting tool. Those features exposed more creator content to the open web, which expanded the surface area that scrapers could reach. That forced a harder response.
For anyone running a business that produces content, the practical question is straightforward. If your company publishes reports, guides, case studies, product documentation, or any other written work on the open web, AI training bots are likely reading it. Whether that bothers you depends on your business, but the tools to control it now exist and are becoming easier to use. Cloudflare offers bot blocking to website operators, including those on free plans. The harder question is whether blocking training bots costs you visibility in AI-generated answers, since some of those answers do send referral traffic back to sources. That trade-off will look different for every organisation.
The broader pattern is straightforward: a norm that served the internet for thirty years, open access in exchange for referral traffic, has broken down. Platforms with large, valuable content libraries are using that as leverage. The ones moving fastest are the ones with the clearest financial incentive to do so.