> ## Content Index
> Fetch the complete content index at: https://nextwith.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Unsealed Times filings add new details on AI training data and publisher harm
- URL: https://nextwith.ai/unsealed-times-filings-add-new-details-on-ai-training-data-and-publisher-harm/
- Published: 2026-09-18T15:06:27.000Z
- Updated: 2026-09-19T05:05:26.000Z
- Description: Unsealed filings in the New York Times v. OpenAI/Microsoft case add new detail on scraped news data, paywall workarounds and alleged publisher harm, while Microsoft and OpenAI dispute how the internal material should be read.
- Author: NextWith.ai Editorial Desk
- Tags: Business, News

Unsealed filings in the New York Times' copyright case against OpenAI and Microsoft became public on Sept. 17, 2026, as the U.S. District Court for the Southern District of New York weighs summary judgment motions. The newly public material adds detail to a dispute over how the companies trained AI systems on news content, how they viewed paywalled journalism and whether their products were cutting into publishers' traffic. [The New York Times reported](https://monorepo-sample1.nyt.net/2026/09/17/technology/microsoft-openai-publishing-industry.html?ref=nextwith.ai) the release of the filings; [TechCrunch](https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal?ref=nextwith.ai) summarized additional claims from the plaintiffs' brief.

## What the newly public record adds

According to The Times, internal Microsoft and OpenAI messages show employees debating whether scraping large volumes of news articles amounted to a serious appropriation of publishers' work, while OpenAI staff warned that chatbots could become increasingly substitutive for news. The reporting says those concerns also extended to the economics of the news business, with one Microsoft document describing a risk that AI products could damage the supply chain they depend on. Those descriptions come from the unsealed material as summarized by the reporters; they do not independently prove the companies' legal liability.

The same reporting says plaintiffs allege OpenAI discussed ways to bypass publishers' paywalls, and that Microsoft chief executive Satya Nadella testified paywalled material should be licensed if a company wants to use it for grounding or training. Microsoft said the internal memos cited in the case did not reflect the company's views, and that Nadella's testimony was about broader shifts in how people consume information rather than a copyright conclusion. OpenAI did not respond to requests for comment in the reporting cited here.

TechCrunch added more scale to the plaintiffs' claims. Citing the filing, it said OpenAI's mid-training datasets contained more than 91,692 copies of works published by the Times, the New York Daily News and the Center for Investigative Reporting, and that a Common Crawl-derived dataset included more than 2 million documents from nytimes.com. TechCrunch also noted that much of the fresh material comes from the plaintiffs' brief, while some underlying exhibits remain sealed.

## Why the filing matters for the case

The practical importance is not that the documents settle the copyright question. It is that they give the publishers more material to argue that the companies knew their systems could displace the original source material and still used the content anyway. That goes directly to the fair-use fight now before Judge Sidney H. Stein, who must decide whether enough of the case survives summary judgment to proceed further.

The next signal to watch is the court's summary judgment ruling and any further unsealing of exhibits. If the judge keeps expanding the public record, the dispute over what was copied, how it was used and how the companies assessed harm to publishers will become clearer before trial.

For teams procuring AI systems, the useful next check is whether a supplier can document the provenance and licensing position of material used for grounding or training. That does not decide this lawsuit, but it makes the dispute a question that can be asked before a tool enters an important workflow.