AI Scraping Threatens the Open Web and Fuels a Data Arms Race
The internet’s foundational structure is facing an unprecedented challenge as artificial intelligence (AI) companies increasingly rely on large-scale web scraping to fuel their large language models (LLMs). This practice, whereas essential for AI development, is creating a “war of attrition” for websites, particularly those reliant on open access and public service, and raising concerns about the future of the open web.
The Rise of AI Scraping and its Impact
In recent months, numerous websites, including openDemocracy and OpenStreetMap, have experienced repeated disruptions due to massive bot traffic. These bots, deployed by AI companies, are designed to scrape data for training LLMs. The scale and sophistication of these scrapers are making it increasingly difficult for website operators to defend against them. Grant Slater, a core developer at OpenStreetMap, described the situation as being “at constant war” with scrapers, noting a “dramatic shift” in their prevalence over the past two years.
These scrapers often bypass standard “robots.txt” protocols, targeting vulnerable website components and consuming significant bandwidth and computing resources. This impacts not only website performance but also the ability of organizations to serve their users effectively. The financial burden of defending against these attacks falls disproportionately on publishers, while AI companies reap the benefits of the extracted data.
Residential Proxies and the Illusion of Legitimate Traffic
A key tactic employed by AI scrapers is the use of residential proxy networks. These networks route internet traffic through the IP addresses of real homeowners, making it difficult to distinguish between legitimate users and automated bots. This allows scrapers to “hide in plain sight,” rotating identities and extracting data at scale. As Slater explains, this creates a situation where “keeping the site usable becomes a constant battle.”
The Threat to an Open and Accessible Internet
Audrey Hingle, a researcher studying bot activity, warns that this trend risks “accelerating enclosure”: more gated content, limited access, and a web that is harder to participate in. Websites may be forced to implement stricter access controls, such as requiring user logins, to combat scraping, potentially hindering the open exchange of information.
Data Control and the Emergence of Data Unions
Concerns over data privacy and control have led to initiatives like the First International Data Union (FIDU), founded by Tony Curzon Price. FIDU aims to empower individuals to use LLMs without surrendering their data to large AI companies. “Our data…is the one chokepoint that ordinary citizens have over platform and BigAI power,” Curzon Price stated. This reflects a growing movement to reclaim control over personal data in the age of AI.
Broader Concerns Surrounding AI Development
The competition to develop advanced AI systems is also raising broader ethical and societal concerns. These include the environmental impact of energy-intensive data centers, the potential for misuse of AI technologies (such as in sexual abuse), and the risk of job displacement. Even within the AI industry, there is a tension between innovation and responsible development, as companies navigate the pressures of both commercial competition and national security interests.
Staying Informed and Supporting Trustworthy Sources
As AI reshapes the information landscape, it is crucial for individuals to be discerning consumers of news and information. This includes scrutinizing sources, consolidating direct links with reputable organizations (through email newsletters, for example), and being aware of the potential for bias and misinformation in AI-generated content. Maintaining connections with trusted media outlets, independent of corporate platforms, is essential for preserving a well-informed public.
Keep reading