Historia
julio 14, 2026
NYT Says OpenAI Hid Key Copyright Evidence as AI Firm Cites User Privacy Defense
In a new court filing, The New York Times and other news organizations have accused OpenAI of withholding evidence in their ongoing copyright infringement lawsuit. The publishers allege that OpenAI misled the court about its ability to search user chat logs and training data for copyrighted material and are seeking sanctions against the company.
The long‑running copyright battle between The New York Times and OpenAI has escalated into a fight over transparency, with publishers accusing the AI company of concealing evidence while OpenAI insists it is protecting user privacy and fair use.
Early lawsuit and discovery battles
The dispute stems from a two‑year lawsuit in which The New York Times alleges OpenAI unlawfully used its journalism to train ChatGPT and then reproduced that work in responses to users.1 News organizations sought access to training data and large samples of chat logs to test how often their copyrighted articles appeared in outputs.2
From the outset, OpenAI argued it lacked the technical ability to search its massive training corpus and that producing extensive ChatGPT logs would be burdensome and raise “user‑privacy concerns.”12
April deposition shifts the case
According to the publishers, an April court‑ordered deposition of OpenAI privacy engineer Vincent (Vinnie) Monaco upended that narrative. Monaco allegedly revealed that OpenAI had already conducted internal searches of its training data for copyrighted journalism and had amassed a database of about 78 million de‑identified ChatGPT conversations to assess potential infringement.1
He also described tools under “Project Giraffe,” including a “Bloom” filter to detect and record regurgitation of text in ChatGPT outputs, implemented shortly after the NYT lawsuit was filed.1
Sanctions motion and dueling narratives
On July 9, The New York Times and other publishers filed a motion accusing OpenAI of “withholding evidence” and failing to share information on “how the company’s A.I. systems are trained and used,” seeking legal sanctions.3 They allege OpenAI deleted billions of ChatGPT outputs after the suit was filed and provided a heavily redacted 20‑million‑log sample the court deemed “unusable.”12
News organizations say this shows OpenAI “repeatedly lying for years to conceal evidence of infringement” and faking an inability to search training data and logs, warranting “serious sanctions.”2
OpenAI counters that the sanctions push is a late‑stage bid to “snoop” through user conversations, portraying it as an effort to “invade the privacy of people who have nothing to do with this case.” The company argues that as the Times has dropped some claims, it is the publishers’ case that is weakening, and vows to defend “users’ privacy and the long‑established principles of fair use.”2
The court must now decide whether OpenAI’s conduct amounts to discovery abuse—or a privacy‑driven stance in a landmark AI copyright fight.