Concerns Over Content Theft and AI Scraping
In recent years, a significant issue has emerged regarding artificial intelligence and its relationship with copyrighted content. Many individuals and organizations have voiced concerns that AI technologies are unlawfully using their content without permission. This problem has become so pronounced that even robust services like Cloudflare, which we rely on and pay substantial fees for, are struggling against sophisticated site-scraping attempts. As a result, we occasionally have to implement captchas to thwart this theft. Without such measures, the performance of our site deteriorates, leading to timeouts that hinder our ability to post new articles or manage user comments.
Cloudflare’s Response to the Crisis
Cloudflare seems to share our frustrations. According to an article in Adweek, the situation has become dire for many publishers:
The nuclear option is gaining traction as web traffic collapses and Google refuses to negotiate with content creators.
For decades, publishers have strived both legally and through gray tactics to boost their rankings in Google Search. However, some influential players in the digital media landscape are now considering an unprecedented move: removing themselves entirely from Google Search.
Recently, Cloudflare issued a firm ultimatum to Google. Starting September 15, new websites joining Cloudflare, as well as customers utilizing the free tier, will have default settings that block “multi-purpose crawlers” from accessing any webpage that contains advertisements. This means crawlers that scrape content for both search indexing and AI training will be denied access, unless site owners choose otherwise.
“We’ve been clear about what we want,” stated Cloudflare Chief Strategy Officer Stephanie Cohen. “We want a technical solution that allows you to be discoverable without having to give your content away for free.”
While some crawlers, like those from Apple and Bing, fall under this categorization, the main target of this initiative is Google, which notoriously employs a single crawler for both website indexing and AI model training.
To be fair, Google has introduced an option known as Google Extended, permitting publishers to opt out of AI training without disappearing from search results. Nevertheless, many publishers remain skeptical, fearing that participation in this program could adversely affect their visibility in search results, as indicated by executives from two media companies.
Challenges with Google for Independent Publishers
What the article does not highlight is Google’s longstanding hostility towards independent publishers, a trend that can be traced back to at least 2014. Websites like The Intercept and Truthout, along with ours, have faced considerable downgrades in search rankings. As such, we find Google’s capacity to help us reach new readers questionable. However, alternatives like Kagi are proving to be much more effective, especially for site-specific searches. With Kagi, users not only gain better search outcomes but also enhanced privacy, as they don’t retain search histories.
To reinforce our skepticism regarding Google’s approach:
Google pays Reddit millions for content access.
Your website’s content is an asset too.
Don’t let every AI crawler scrape it for free. Control who can access your content while keeping Google Search operational.
Cloudflare AI Crawl Control + Pay Per Crawl are just the beginning. pic.twitter.com/WRxXzEx6KD
— Cuong Thach (@cuongthach_) July 10, 2026
Industry Moves Towards Control Over Content
Notably, several prominent organizations are aligning themselves with Cloudflare’s initiatives. As reported by Adweek:
USA Today Inc., the parent company of USA Today and numerous regional news sites, is considering its options, according to CEO Mike Reed. The company has been addressing declines in search traffic by diversifying its audience through newsletters, social media, and events. It has maintained stable traffic levels, consistently achieving over 1 billion pageviews monthly for the past three years.
Looking ahead, USA Today Inc. plans to derive its monetization strategy related to AI from licensing agreements with companies like Meta, Microsoft, and Amazon. In contrast, Google has not established any such agreements with publishers.
Consequently, USA Today Inc. is prepared to withdraw from Google within the next six to twelve months, as mentioned by Reed. Additionally, Beehiiv, a creator network, has recently partnered with Cloudflare, enabling its creators to block the Google crawler.
“I wouldn’t consider this a monumental decision since we are already blocking other crawlers,” Reed explained. “Those with licensing agreements will access our content; those without will be blocked.”
Although USA Today Inc. and Beehiiv are the first major entities to undertake this approach, executives from numerous media companies are contemplating similar measures against the Google Bot, according to an anonymous source involved with Google negotiations.
The potential revenue from AI scraping is significant. For example, I am expecting $3,000 from a class action settlement with Anthropic. One can only imagine how much those exploiting the extensive body of work at Naked Capitalism should be compensating its creators.
The Broader Impacts of AI on Original Content
In summary, the developments surrounding AI present serious challenges for original content creators, as it often leads to the generation of derivative works that lack depth and value. While informational shortcuts are appealing to some, the consequences for critical thinking and comprehensive understanding are concerning.