Unsealed NYT filing: Microsoft, OpenAI staff called AI scraping 'theft'
Original source
Microsoft exec called AI scraping 'the largest theft of labor in human history'
Hacker News →Newly unredacted filings in The New York Times’ three-year-old copyright suit against OpenAI and Microsoft quote the companies’ own employees describing their AI training practices in damning terms. A Microsoft director of applied science, Brent Hecht, reportedly called the mass ingestion of published work ‘an astonishing theft of unprecedented proportions’ and ‘the largest theft of labor in human history,’ while OpenAI’s head of ChatGPT warned that the product posed an ‘existential threat’ to publishers because it is ‘largely substitutive’ of their work. The filings also allege the companies bypassed paywalls without detection, assembled training sets like WebText and Project Mango that leaned heavily on news content, and deliberately stripped copyright notices so models wouldn’t reproduce them.
The admissions matter because they cut directly against the fair-use defense that has so far served AI firms well in court. Fair use turns partly on whether a copy harms the market for the original, and internal Microsoft data cited in the brief shows its Copilot answer engine drove click-through rates to nytimes.com down as much as 93% versus traditional Bing search—a ‘doom loop’ one presentation warned would degrade both the models and the open web. Satya Nadella testified that paywalled content should be licensed for training, and both he and OpenAI leadership conceded on the record that chatbots substitute for visiting the source, undermining the argument that training is transformative rather than competitive.
The scale is notable: OpenAI mid-training datasets allegedly held more than 91,000 copies of works from the Times, Daily News, and Center for Investigative Reporting, with a Common Crawl-derived set pulling over two million documents from nytimes.com alone. Important caveat—most of these quotes come from the Times’ own brief rather than the still-sealed exhibits, so they arrive stripped of original context. Courts and even the Trump administration have leaned toward AI companies’ fair-use arguments, but these internal statements hand publishers concrete evidence that the firms understood the harm they were causing. Microsoft and OpenAI declined to comment.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.