The New York Times, New York Daily News, and other media outlets accused OpenAI of lying to the court about its ability to search ChatGPT logs and training data for copyrighted works. They contend OpenAI searched its training data and chat logs for patented news articles before lawsuits were filed, contradicting claims that such searches were infeasible or burdensome [1, 2, 3, 4].

The newspapers allege OpenAI deleted billions of ChatGPT conversation logs or rendered them unsearchable after the lawsuits began, violating court preservation orders [1, 3]. The plaintiffs are asking a federal court in Manhattan to sanction OpenAI for discovery misconduct including hiding, destroying evidence, and making misrepresentations about its ability to search data [1, 4]. Ian Crosby, lead attorney for the New York Times, said, “For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court” [1]. Steven Lieberman, attorney for the New York Daily News, added, "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism" [4].

OpenAI denies wrongdoing and disputes the allegations. It argues turning over chat logs would violate user privacy and claims the news organizations’ lawsuit is weakening as they drop some claims. An OpenAI spokesperson said, "As the Times’ case weakens and they’ve been forced to drop claims against us, they’re persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations" [2, 4]. Another representative said the company will "continue defending our users’ privacy and the long-established principles of fair use" [2].

Internal OpenAI documents revealed the company maintained a database of about 78 million de-identified ChatGPT conversations to evaluate potential copyright infringement [3]. It used tools such as a “Bloom” filter and “Project Giraffe” to detect reproduction of copyrighted content in AI outputs [3]. In December 2025, OpenAI submitted a sample of 20 million chat logs to the court, heavily redacted and much smaller than the 120 million requested by plaintiffs, limiting their usability [3].

The dispute began in 2023 when the New York Times sued OpenAI and Microsoft for alleged unauthorized use of news articles in AI training [1, 4]. In April 2026, OpenAI privacy engineer Vincent Monaco was re-deposed and revealed the company had conducted searches for copyrighted content, contradicting earlier claims [2, 3]. On July 9, 2026, the newspapers filed a motion seeking court sanctions against OpenAI for evidence concealment and deletion, escalating the copyright battle [1, 4].

The court will now consider the plaintiffs’ motion for sanctions as the case moves forward.