Publishers Limit Internet Archive Access Amid AI Scraping Concerns
A growing number of news publishers are restricting access to their content for the Internet Archive, the organization behind the Wayback Machine, due to concerns about artificial intelligence (AI) companies scraping their articles for training data. This move risks diminishing a crucial historical record of online journalism.
The Internet Archive and the Wayback Machine: A Digital Record
For nearly three decades, the Internet Archive has been preserving web pages, including those of news organizations, through its Wayback Machine. This archive contains over one trillion archived web pages and serves as a vital resource for journalists, researchers, and legal proceedings. The Wayback Machine often provides the only reliable record of how news stories were originally published, as articles are frequently edited or removed over time. Internet Archive
Publishers Push Back Against Scraping
The Guardian was among the first major news organizations to limit the Internet Archive’s access to its published articles, minimizing the potential for AI companies to scrape content via the nonprofit’s repository. The publisher achieved this by excluding itself from the Internet Archive’s APIs and filtering article pages from the Wayback Machine’s URLs interface. Regional homepages, topic pages, and other landing pages remain accessible in the archive. Nieman Lab
The Modern York Times has also taken steps to block the Archive from crawling its website, employing technical measures beyond standard robots.txt rules. This action raises concerns about the loss of a historical record that has been relied upon for decades.
AI Training and Copyright Concerns
Publishers are increasingly concerned about AI companies using their copyrighted material to train AI models without permission. Several publishers, including The New York Times, are currently pursuing legal action against AI companies over this issue, questioning whether such training constitutes fair use. Nieman Lab
Fair Use and the Role of Archives
Despite the legal disputes, experts argue that archiving and search are established forms of fair use. Courts have previously recognized that creating searchable indexes requires copying underlying material, as demonstrated in cases involving Google Books. The Internet Archive operates on a similar principle, preserving the web’s historical record for research and public access. Just as physical libraries preserve newspapers, the Internet Archive preserves the digital record.
According to Internet Archive staff, Wikipedia alone links to over 2.6 million news articles preserved at the Archive, spanning 249 languages, highlighting its widespread use and importance. Internet Archive
The Risk of a Vanishing Historical Record
By blocking the Internet Archive, publishers risk erasing a significant portion of the web’s historical record. Whereas disputes over AI training require resolution, sacrificing the public record in the process would be a detrimental and potentially irreversible mistake. The legal principles protecting search engines and web archiving are already well-established, even if courts impose limits on AI training.
Keep reading