Seattle Times and Newsday filed a lawsuit in the U.S. District Court for the Southern District of New York on September 4, accusing OpenAI and Microsoft of using their copyrighted news content without authorization to train AI models including ChatGPT, Copilot, and Bing AI [1, 2, 3, 4, 5]. The lawsuit alleges the companies collected content from the news organizations’ websites, including paywalled articles, and incorporated it into datasets that power their AI products [1, 2, 3, 4, 5].
The complaints claim these AI models can reproduce news stories, rewrite articles in similar ways, and answer user questions based on the content, which reduces traffic and subscription revenue for the publishers [1, 2, 3, 4, 5]. The lawsuit demands the court order OpenAI and Microsoft to destroy copies of their news content and any datasets or AI models containing it [1, 2, 3, 5].
Seattle Times and Newsday also allege OpenAI and Microsoft bypassed paywalls during their data scraping process to obtain restricted content [4]. The lawsuit notes that Microsoft and OpenAI have previously funded some of Seattle Times’ journalism projects and fellowships [6].
Microsoft called the lawsuit unexpected but emphasized its understanding of the importance of local news. A Microsoft spokesperson said the company is "always happy to sit down and explore solutions to this type of dispute" [1]. OpenAI stated its AI models are trained on publicly available data and claim that their use is protected under fair use, declining further comment [1, 2, 3, 5].
This legal action follows a 2023 lawsuit by The New York Times, which accused OpenAI and Microsoft of using millions of its news articles without permission to train AI models; that case remains ongoing [1, 2, 3, 4, 5]. The New York Times spokesperson said, "AI and creators can coexist, AI companies just need to pay reasonable fees for content" [5]. Meanwhile, the U.S. Department of Justice, not party to these suits, supports the view that using large internet text datasets for AI training falls under fair use [5].
Dozens of similar copyright lawsuits have been filed against AI companies like OpenAI, Anthropic, and Meta over unauthorized use of copyrighted materials to train large language models [1, 2, 3, 5].
The Seattle Times and Newsday lawsuits mark the latest legal challenges as courts address the boundaries of copyright in AI training data. No court dates or further proceedings have been announced yet.