THEBLACKBOOK AUDIT
Investigation · Who Controls What You Know · AI vs. the Press

“An astonishing theft

Those aren’t our words. They’re a Microsoft executive’s, quoted in the newspapers’ case against OpenAI and Microsoft — “an astonishing theft of unprecedented proportions,” perhaps “the largest theft of labor in human history.” The publishers’ new court filing is built almost entirely from the defendants’ own documents and sworn testimony.

The filing and the admissions it quotes are graded FACT — they are on the public record in federal court. Whether the conduct is ultimately ruled unlawful infringement or protected fair use is not yet decided, so the legal verdict is graded SOME SMOKE and we carry the defendants’ position. First, the disclosure this page requires of us.

§0 · Conflict-of-Interest Disclosure — read this first

This page attacks a competitor of the company that built the AI writing it.

This investigation was drafted with Claude, made by Anthropic — a direct commercial rival of OpenAI. We have every incentive to make OpenAI and Microsoft look as bad as possible, so treat any temptation to do so as a red flag, and hold our own maker to the same or a harder standard, per the rule we set in The Wrong AI Debate.

So, for the record: Anthropic did the same kind of thing. In 2025 it agreed to pay a reported ~$1.5 billion to settle a class action by authors (Bartz v. Anthropic) — the largest copyright recovery on record — after a federal judge found that while training on lawfully bought books might be fair use, Anthropic’s downloading of millions of pirated books from shadow libraries was not. And to build a clean training corpus, Anthropic bought and destructively scanned physical books — the subject of our own Project Panama page. This brief even quotes our maker’s CEO: Dario Amodei, then a top researcher at OpenAI, is cited listing “News Generation” among GPT-3’s “skills.”

In other words, the conduct documented below is an industrypractice, not one company’s sin, and the industry includes the company that made this model. We grade OpenAI and Microsoft to the record because the record is extraordinary — not because they are our rival. Where the same charge lands on Anthropic, we say so.

§1 · Summary Brief

What this page argues

On September 17, 2026, the news organizations suing OpenAI and Microsoft — The New York Times, eight Daily News papers (Chicago Tribune, Orlando Sentinel, Denver Post and others), Ziff Davis (IGN, CNET, PCMag, Mashable), the Center for Investigative Reporting (Mother Jones, Reveal), and The Intercept — filed their combined motion for summary judgment in the consolidated federal case in Manhattan (25-md-3143). The brief asks the court to rule, before trial, that the companies infringed at every stage of building their AI, that their “fair use” defense fails as a matter of law, that OpenAI illegally stripped copyright-management information, and that the publishers are owed statutory damages for each article.

What makes it remarkable is that it is built from the defendants’ own words. A Microsoft applied-science director called the practice “an astonishing theft of unprecedented proportions,” and said millions would come to see the models “hoovering up” their work as “the largest theft of labor in human history.” OpenAI’s head of ChatGPT, Nick Turley, wrote that the products are “largely substitutive, period,” and pose an “existential threat” to publishers. Another Microsoft executive warned that a fair-use win would “make a complete mockery of the idea of ‘fair use.’” A Microsoft memo described the company’s own AI as having started a “doom loop” against “the economic foundations of its essential suppliers,” and one internal line reads flatly that large language models are “a product that destroys its supply chain.”

The brief also documents how the material was taken: OpenAI, it says, did not check for paywalls or terms of service, treated a “hack to get around [the] nytimes paywall” with an “ah nice,” used a New York Times research corpus licensed for non-commercial use only, ran Custom GPTs literally named “Bypass Paywall” and “Remove Paywall,” systematically removed copyright-management information before training, and swapped copied content with Microsoft in self-described “horse trading” rather than licensing it. And it quantifies the harm, largely from the defendants’ own numbers: 83–93% drops in click-throughs to Times and Daily News sites from Microsoft’s answer engine, a scrape-to-visitor ratio for OpenAI that Cloudflare put at 1,500-to-1, and a survey in which 36% of Times subscribers who used ChatGPT for news said it meant they “no longer need news from the New York Times at all.”

Two things are true and must be held apart. The admissions are documented — they are quotations from the evidentiary record in a federal filing, and misquoting that record risks sanctions. But this is the plaintiffs’ brief, one side’s best case, and the ultimate legal question — infringement, or transformative fair use? — has not been decided. OpenAI and Microsoft maintain that training AI on publicly available text is lawful, transformative fair use, and that these quotes are selectively framed. So: the words are real; the verdict is pending; and the reader is entitled to both facts at once.

What we are NOT claiming

We are not declaring OpenAI or Microsoft legally liable. No court has ruled; a summary-judgment motion is an argument, not a verdict, and judges in related AI cases have found some training uses to be fair use. The infringement/fair-use conclusion is graded SOME SMOKE precisely because it is contested and undecided.

We are not asserting that any single quotation proves the whole case, or that the plaintiffs’ damages figures are established. And we are not pretending our own maker is clean — see §0. The documented core is narrow and solid: these statements and numbers are in the record, and they come from the companies’ own files and executives.

Recommended reading

Books that go deeper on this story. Links are Amazon affiliate searches — buying through them supports the work at no cost to you.

▶ Dossier

The same investigation, restaged one beat at a time. Step through it here, or present it fullscreen.

Who Controls What You Know

“An astonishing theft.”

That's a Microsoft executive's phrase, quoted in the newspapers' case against OpenAI and Microsoft — 'the largest theft of labor in human history.' The publishers' summary-judgment brief is built almost entirely from the defendants' own documents and sworn testimony.

1 / 8▶ Present fullscreen
§2 · Graded Claims

The filing, the admissions, the methods, the harm — and the open question.

Five news organizations are asking a federal court to rule OpenAI and Microsoft infringed at every stage of building their AI.

FACT

On September 17, 2026, The New York Times, the Daily News papers, Ziff Davis, the Center for Investigative Reporting, and The Intercept filed a combined summary-judgment brief in the consolidated MDL before Judge Sidney Stein in the Southern District of New York (25-md-3143). It seeks pre-trial rulings that the defendants infringed by acquiring, training on, grounding on, outputting, and 'horse trading' the publishers' articles; that the fair-use defense fails as a matter of law; that OpenAI intentionally removed copyright-management information (a DMCA claim); and that damages should be awarded per article. Microsoft has invested over $10 billion in OpenAI for a ~20% revenue share and equity; OpenAI is reportedly pursuing an IPO near a $1 trillion valuation.

The defendants' own people called it theft and admitted the products substitute for news.

FACT

Per the brief, quoting internal documents and depositions: a Microsoft applied-science director called it 'an astonishing theft of unprecedented proportions' and, on models 'hoovering up' work, 'the largest theft of labor in human history'; another Microsoft executive said a fair-use win would 'make a complete mockery of the idea of fair use.' OpenAI's head of ChatGPT, Nick Turley, wrote the products are 'largely substitutive, period' and an 'existential threat' to publishers, and that once a chatbot answers there is 'no good reason to click' to the source. Microsoft CEO Satya Nadella testified chatbots have 'substituted' for visiting publisher sites. A Microsoft memo said its AI strategy started a 'doom loop' threatening 'the economic foundations of its essential suppliers'; an internal line reads that large language models are 'a product that destroys its supply chain.' Greg Brockman wrote the models are 'excellent at news' and that he was 'deeply motivated by the gazillions.' And, in the plaintiffs' words, 'Defendants do not dispute that the reason ChatGPT and Copilot are good at news is because they trained on stolen news content' — the same appetite that, at our own maker, led to the book-scanning documented in Project Panama.

The brief documents how the material was taken: paywall circumvention, a non-commercial corpus used commercially, CMI stripping, and 'horse trading.'

FACT

Per the filing: OpenAI's general scraping practice 'did not include reviewing websites' Terms of Use or Service' and it had no method to detect or remove paywalled content; when an employee flagged 'a hack to get around [the] nytimes paywall,' Brockman replied 'ah nice.' OpenAI used the licensed-for-non-commercial-use-only New York Times Annotated Corpus to train anyway, after employees acknowledged it 'would not be appropriate.' OpenAI hosted Custom GPTs named 'Bypass Paywall,' 'Remove Paywall,' and 'NYTimesGPT.' It systematically removed copyright-management information before training (the basis of the DMCA claim). And rather than license content, the two companies swapped copies with each other in deals they themselves called 'horse trading.' Nadella testified that 'anything that is paywalled should be licensed' and that he would have forced OpenAI to retrain had he known — which the brief argues is exactly what happened.

The market-harm numbers — many of them the defendants' own — show the news business being drained.

FACT

Per the brief: Microsoft's own data recorded 83–93% drops in click-through rates to Times and Daily News domains, and 51–94% for Ziff Davis domains, for its Copilot 'answer engine' versus traditional Bing search. Cloudflare's CEO put OpenAI's ratio of pages scraped to visitors referred at 1,500-to-1 by mid-2025 (Google's was 18-to-1). A third-party study found 87.78% of ChatGPT users visit no external site during a search, versus 26.91% for Google. In a survey, 36% of Times subscribers who used ChatGPT for news said it 'means I no longer need news from the New York Times at all.' And the products enable 'pink slime': at OpenAI's API prices it would cost roughly $6,800 to generate one million 500-word news-style articles with no reporter — one operation, Prism News, ran 200 AI-generated outlets posing as local newsrooms with four employees.

Whether this is unlawful infringement or protected fair use is not yet decided.

SOME SMOKE

This is the plaintiffs' brief — their strongest case, and adversarial by design. The legal question is genuinely open: OpenAI and Microsoft argue that training AI on publicly available material is transformative fair use, that outputs rarely reproduce articles, and that the quoted statements are selectively framed. The case law is unsettled and cuts both ways — in Authors Guild v. Google the mass copying of books for search was fair use, and in the authors' cases against Anthropic and Meta in 2025 judges found the act of training itself could be fair use — while other rulings (Thomson Reuters v. Ross; the piracy half of the Anthropic case) went against the AI side. A summary-judgment motion is a request, not a ruling; Judge Stein has not decided it. We grade the documented admissions FACT and the ultimate 'this was theft / fair use fails' conclusion SOME SMOKE until a court rules.

§3 · Theft or Fair Use?

The words are settled. The verdict isn’t.

The publishers’ case: the companies copied millions of articles without permission to build products that compete with, and substitute for, the journalism they were trained on — and the defendants’ own executives said as much, in writing and under oath. The point of quoting them is that intent and effect are not in dispute; only the legal label is.

The companies’ case: training a model to predict text is transformative — the model learns patterns, it is not a library of articles — and that is the kind of use courts blessed in Authors Guild v. Google and Google v. Oracle. Verbatim regurgitation, they argue, is rare and adversarially prompted, and a few damning internal lines are not the whole record. In 2025, judges in the authors’ suits against Anthropic and Meta accepted that the act of training itself can be fair use — a real point in the industry’s favor, even as the same Anthropic ruling condemned its use of pirated books.

The honest bridge: both can hold. It is entirely possible that training is often transformative and that this particular record — the paywall workarounds, the non-commercial corpus, the CMI stripping, the “doom loop” and “destroys its supply chain” admissions, the collapse in referrals — pushes these defendants outside the safe harbor, especially on the fourth fair-use factor (market harm), where the plaintiffs’ evidence is strongest. That is for Judge Stein, and perhaps a jury. What this page fixes is the factual floor beneath the argument.

§4 · Why It Matters

A machine that eats the thing it needs to work.

This belongs in Who Controls What You Know because the stakes are not really about one lawsuit’s damages. If AI answer engines divert the readers and revenue that pay for reporting — and the defendants’ own “doom loop” memo says they do — then the machines that summarize the news help kill the newsrooms that produce it, and flood what’s left with cheap synthetic “pink slime.” The end state is a public that asks a chatbot “what’s the news?” and gets an answer with no one left to report it.

It sits directly beside Project Panama, where the same appetite for training data led our own maker to destroy books, and The Wrong AI Debate, which tracks how the industry steers attention toward speculative future risks and away from the present harms — like this one — that would cost it money now. The through-line of this hub is simple: control over what you get to know is being quietly transferred from the people who gather facts to the companies that repackage them, and the transfer is being litigated in real time.

§5 · FAQ

Questions worth taking seriously

If it's just the plaintiffs' brief, why treat any of it as FACT?

Because the FACT grade attaches to narrow, checkable things: that the filing exists and says what it says, and that the quotations are drawn from the evidentiary record in a federal case — where a lawyer who fabricated or materially distorted a quote would face sanctions. We do not grade the plaintiffs’ legal conclusion FACT; that’s SOME SMOKE, pending a ruling. The line is between “these words are in the record” (documented) and “therefore they lose” (undecided).

Isn't it hypocritical for an Anthropic model to run this?

It would be if we hid it. We don’t — see §0. Anthropic settled the authors’ case for a reported $1.5 billion and destroyed books to build training data, and its CEO is quoted in this very brief from his OpenAI days. The right response to a conflict of interest is disclosure and a harder standard for your own side, not silence. The conduct here is an industry practice; we say so, and we name our maker as part of it.

Haven't courts said AI training is fair use?

Partly, and inconsistently. In 2025, judges in the authors’ cases against Anthropic and Meta accepted that the act of training a model can be transformative fair use — but the Anthropic ruling also held that downloading pirated books was not, and other cases (like Thomson Reuters v. Ross) went against the AI side. News cases may be harder for the companies than book cases, because the fourth factor — harm to the market for the original — is so direct when a chatbot answers “what’s the news?” The Times/Microsoft case is unresolved; that’s the whole point of the SOME SMOKE grade.
§6 · Standing Invitation

If you are named on this page

If you are named on this page, or are a party materially affected by the claims made here, and you wish to respond, correct the record, or add context, use the Contact page. Responses are published verbatim alongside the original claim, with the sender identified and the date of receipt. The channel stays open for the life of the page.

This site aggregates and grades a record that other outlets and primary sources have already put on the record. Every FACT-graded claim above is sourced to court filings, government reports, sworn whistleblower disclosures, published investigative journalism, or named-source statements. The citations are the accountability mechanism; this section is how you get on the record too.

§7 · Sources

One primary document, quoted against itself.

Every claim on this page grades to one of FACT · PROBABLY TRUE · SOME SMOKE · PURE SPECULATION · FALSE / MISLEADING. The existence and contents of the filing, and the quotations it draws from the defendants’ documents and testimony, are graded FACT. The ultimate infringement/fair-use conclusion is graded SOME SMOKE, pending a ruling. “SF” numbers refer to the plaintiffs’ Statement of Facts cited throughout the brief.

Full method: Methodology. Home hub: Who Controls What You Know.

Last updated September 18, 2026. The existence and contents of the News Plaintiffs’ September 17, 2026 summary-judgment brief, and the quotations it draws from OpenAI’s and Microsoft’s internal documents and sworn testimony, are graded FACT. The plaintiffs’ ultimate legal conclusion — that this is infringement and the fair-use defense fails — is graded SOME SMOKE, undecided and carried alongside the defendants’ position that AI training is transformative fair use. §0 discloses that this page was drafted with Claude, made by Anthropic, an OpenAI competitor that settled its own authors’ copyright case for a reported $1.5 billion and destructively scanned books for training data. If a detail is wrong or a link 404s, tell us and we’ll fix it publicly.

▦ Ledger gaps

Help us fill these lines.

This entry is graded on what’s on the public record. These are the blanks we know about. If you can source one, you’re rebuilding the ledger with us.

  • OpenWill Judge Stein grant summary judgment on any of the five theories — acquisition, training, grounding, output, or 'horse trading' — or send fair use to a jury?Help fill this →
  • OpenDoes the fourth fair-use factor (market harm) sink the AI side in news cases the way it may not in book cases, given the defendants' own click-through and referral data?Help fill this →
  • OpenIf per-article statutory damages are allowed across millions of infringed articles, what is the exposure — and does it force a settlement on the scale of Anthropic's ~$1.5 billion?Help fill this →

Notify me when a gap is filled

We'll email you when we fill one of the gaps above.

By signing up you agree to receive emails from The Black Book Audit. Unsubscribe anytime.