Britannica and Merriam-Webster Sue OpenAI Over Copyright Infringement
Encyclopedia Britannica, alongside its subsidiary Merriam-Webster, has filed a lawsuit against OpenAI alleging “massive copyright infringement.” The publisher claims OpenAI scraped nearly 100,000 of its online articles to train its large language models (LLMs) without permission [1], [2]. The lawsuit expands beyond training data, addressing concerns about content reproduction and trademark violations.
Allegations Against OpenAI
Britannica accuses OpenAI of multiple infringements, including:
- Unauthorized Use of Copyrighted Material: OpenAI allegedly used nearly 100,000 copyrighted articles to train its LLMs without authorization [2].
- Verbatim Reproduction of Content: The lawsuit claims ChatGPT generates responses containing “full or partial verbatim reproductions” of Britannica content [2], [3].
- RAG Workflow Infringement: Britannica contends that OpenAI’s retrieval augmented generation (RAG) system, which scans the web for updated information, violates copyright when it uses Britannica articles [2].
- Trademark Violation: The publisher alleges OpenAI violates the Lanham Act when ChatGPT “hallucinates” false information and attributes it to Britannica [2], [4].
Impact on Revenue and Trust
Britannica argues that ChatGPT “starves web publishers…of revenue” by providing direct answers that compete with publisher content [2]. The company asserts that ChatGPT’s inaccuracies jeopardize “the public’s continued access to high-quality and trustworthy online information” [2].
Broader Legal Landscape
Britannica is not alone in pursuing legal action against OpenAI. The New York Times, Ziff Davis (owner of Mashable, CNET, IGN, PC Mag, and others), and numerous newspapers across the U.S. And Canada, including the Chicago Tribune and the Toronto Star, have also filed lawsuits alleging copyright infringement [2]. Britannica also has a pending lawsuit against Perplexity AI [2].
Legal Precedent and Challenges
The legal question of whether using copyrighted content to train LLMs constitutes infringement remains largely unresolved. A recent case involving Anthropic resulted in a $1.5 billion settlement after a judge determined the company illegally downloaded millions of books, even though the use of copyrighted data for training was considered “transformative” [2].
Keep reading