THEBLACKBOOK AUDIT
Investigation · The Tech Right

“Ah nice.”

Told a colleague had built “a hack to get around nytimes paywall,” OpenAI's president answered in two words. A Microsoft executive had another phrase for the same enterprise: “the largest theft of labor in human history.”

In September 2026, a less-redacted court filing in the copyright case The New York Times v. OpenAI and Microsoft put internal messages and memos into the public record. Among them: OpenAI president Greg Brockman replying “ah nice” when an OpenAI researcher described a hack to circumvent the Times's paywall; a Microsoft director calling the training practice “the largest theft of labor in human history”; and evidence that the companies deliberately stripped copyright notices and scraped news content at massive scale. This spoke of The Tech Right examines what those documents show — and, because it is being written by a model made by one of OpenAI's direct competitors, it does so under an unusually sharp conflict of interest, disclosed in full below.

§0 · Conflict of Interest Disclosure

This page attacks OpenAI — and it was drafted by a model made by OpenAI's direct competitor

Black Book Audit uses Claude, built by Anthropic, in its research and drafting. Anthropic is a direct commercial rival of OpenAI. A Claude-drafted piece that makes OpenAI look like a lawbreaking bad actor therefore has an obvious incentive to exaggerate — the mirror image of the conflict we disclose in The Wrong AI Debate, where our maker benefits from the frame we question. We handle it the same way: by grading harder against our own interest, not softer. Concretely, that means this page rests only on OpenAI's and Microsoft's own words — internal messages, memos, and sworn testimony — as quoted in the plaintiffs' court filings and corroborated by multiple independent outlets; we flag that the underlying exhibits remain sealed and that these quotes are presented without their full original context; and we refuse to declare a legal verdict that no court has reached. If anything here reads as a competitor's model eager to convict a rival, that is the failure to watch for — and we invite correction.

§1 · Summary Brief

What this page is about

The New York Times and other publishers sued OpenAI and Microsoft in 2023, alleging that training large language models on their copyrighted articles — often pulled from behind paywalls — was mass infringement. OpenAI's defense is fair use: that learning from text to build a new capability is transformative and lawful. In September 2026, a less-redacted version of the plaintiffs' summary-judgment brief exposed internal communications the companies had fought to keep sealed, and they are unflattering.

According to the filing, when an OpenAI researcher (named as Nick Ryder) described “a hack to get around nytimes paywall,” OpenAI president Greg Brockman replied “ah nice.” A January 2023 internal memo by Microsoft's director of applied science, Brent Hecht, called the scraping “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” The brief describes deliberate stripping of copyright notices from training data, datasets containing tens of thousands of copies of Times articles, and internal admissions that the resulting products are “largely substitutive” for the journalism they were trained on — a Microsoft analysis reportedly found its Copilot cut click-through to the Times's site by as much as 93%. Those substitution admissions matter because fair use turns partly on whether a use harms the market for the original. Whether any of this is ultimately illegal is unresolved: courts have so far leaned toward AI companies on fair use, and the Trump administration filed a brief siding with OpenAI. This page grades what the documents show, and holds the legal conclusion open.

What we are NOT claiming
We are not declaring that OpenAI or its founders “broke the law” — that is the plaintiffs' contention and the argument of the op-ed that prompted this page, but it is unadjudicated, the fair-use defense is live, and no one has been charged. We are not treating the quotes as fully contextualized: they come from the plaintiffs' brief, the underlying exhibits are sealed, and a two-word reply like “ah nice” can read differently with its full thread. We are not resting on a single source: the “ah nice” exchange and the “theft of labor” memo are reported by multiple independent outlets quoting the same filing. And — per §0 — we are not pretending to be a neutral narrator: the model writing this is made by an OpenAI competitor, so we lean on the companies' own words and decline the verdict. What is not seriously in dispute: the filing exists, the quotes are in it, and the substitution and scraping admissions are on the record.
Recommended reading

Books that go deeper on this story. Links are Amazon affiliate searches — buying through them supports the work at no cost to you.

▶ Dossier

The same investigation, restaged one beat at a time. Step through it here, or present it fullscreen.

The Tech Right

“Ah nice.”

Told a colleague had built 'a hack to get around nytimes paywall,' OpenAI's president answered in two words. A Microsoft exec had another phrase for the same enterprise: 'the largest theft of labor in human history.'

1 / 9▶ Present fullscreen
§2 · Thesis

The claim this page defends

OpenAI and Microsoft's own internal records, exposed in the New York Times copyright case, show a culture that treated other people's paywalled, copyrighted work as free raw material — circumventing paywalls, stripping copyright notices, and privately describing the practice as theft — while building products their own staff called substitutes for the journalism they were trained on. Whether that is ultimately ruled illegal is unresolved and we do not prejudge it; what the documents establish is intent and awareness, in the companies' own words.

§3 · Timeline

From a lawsuit to the unsealed messages: 2023 – 2026

  • Dec 2023. The New York Times sues OpenAI and Microsoft, alleging mass copyright infringement in training data.
  • 2024–2025. The case grinds through discovery; OpenAI presses a fair-use defense. Courts elsewhere lean toward AI firms on fair use.
  • Sept 2, 2026. The Trump administration files a brief siding with OpenAI on training AI models on copyrighted material.
  • Sept 17, 2026. A less-redacted version of the publishers' summary-judgment brief is filed, quoting internal messages and memos.
  • Sept 2026. Reporting surfaces the “ah nice” exchange, the “largest theft of labor” memo, and the copyright-stripping and substitution admissions.
  • Ongoing. The core legal question — whether training on copyrighted work is fair use — remains undecided; the exhibits behind the quotes remain sealed.
§4 · Graded Claims

The record, claim by claim

Told of a 'hack to get around nytimes paywall,' Brockman replied 'ah nice.'

FACT

The anchor of the story is a two-word message. According to the plaintiffs' filing, when OpenAI researcher Nick Ryder told OpenAI president Greg Brockman about 'a hack to get around nytimes paywall,' Brockman answered 'ah nice.' The exchange is reported by multiple independent outlets quoting the same filing, so its existence is not in serious doubt. We fence it in one way, per §0: the underlying exhibit is sealed and the quote is presented without its full thread, so we treat it as evidence of a casual, approving attitude toward paywall circumvention at the top of OpenAI — which is what it plainly shows — rather than as a signed confession of a crime.

“[A researcher described] a hack to get around nytimes paywall. — [Greg Brockman:] ah nice. — internal OpenAI messages, as quoted in the News Plaintiffs' summary-judgment brief (Sept 2026)”

A Microsoft director called it 'the largest theft of labor in human history.'

FACT

The most quotable indictment comes from inside Microsoft, not from the plaintiffs. In a January 2023 internal memo, Microsoft's director of applied science, Brent Hecht, described the AI training practice as 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history.' It is a striking phrase precisely because it is a senior employee of one of the defendants characterizing his own side's conduct. We grade the fact that the memo says this as FACT, corroborated across outlets; we do not adopt 'theft' as our own legal conclusion, since whether it is legally theft (versus fair use) is the very question the court has not answered.

“An astonishing theft of unprecedented proportions … the largest theft of labor in human history. — Brent Hecht, Microsoft director of applied science, internal memo (Jan 2023), as quoted in the filing”

The companies deliberately circumvented paywalls and stripped copyright notices.

FACT

The 'ah nice' message was not an isolated quip but part of a described practice. The filing alleges OpenAI employees developed ways to bypass publisher paywalls without detection, and that training data was deliberately processed to strip out copyright-management information, because researchers 'wouldn't want [the] model outputting' copyright notices to users. Removing copyright notices is legally significant in its own right — it implicates a distinct provision of copyright law — and, more to this hub's point, it shows an awareness that the material was copyrighted and a decision to obscure that fact. We grade the presence of these allegations-with-quotes in the filing as fact, while noting they are the plaintiffs' characterization pending the sealed exhibits and any defense rebuttal.

The copying was industrial in scale — tens of thousands of Times works, millions of pages.

FACT

The documents put numbers to the scraping. Per the filing, OpenAI's mid-training datasets alone contained more than 91,692 copies of works published by the Times, the Daily News, and the Center for Investigative Reporting; a Common Crawl-derived dataset included more than 2 million documents from nytimes.com; and a Microsoft-OpenAI dataset ('Project Mango') contained copies of at least 160,903 unique works from the news plaintiffs. These are the plaintiffs' figures, drawn from discovery, and they establish that whatever one calls it legally, the use of this journalism was vast and systematic, not incidental. Scale bears directly on the fair-use question of how much of the original was taken.

Their own staff called the products 'substitutive' — the admission that cuts against fair use.

FACT

The most legally pointed material is the companies' own assessment of market harm. Fair use weighs whether a use substitutes for, and harms the market for, the original. Yet OpenAI's head of ChatGPT, Nick Turley, reportedly wrote that publishers face an 'existential threat' from products like the chatbot, which are 'largely substitutive' and 'will get more and more substitutive as they get better.' A Microsoft analysis found its Copilot 'answer engine' cut click-through to the Times's domain by as much as 93% versus traditional search, which an internal presentation called a 'doom loop.' And CEO Satya Nadella testified that anything paywalled 'should be licensed' for training. These are admissions, in the defendants' own voices, on the exact factor their legal defense most needs to win.

“[Publishers face an] existential threat [from products that are] largely substitutive … and will get more and more substitutive as they get better. — attributed to Nick Turley, OpenAI head of ChatGPT, in internal communication quoted in the filing”

Whether any of it is illegal is unresolved — and the government has sided with OpenAI.

SOME SMOKE

Here we hold the line the op-ed that prompted this page does not. The claim that OpenAI's founders 'broke the law' is a serious argument with real supporting evidence — the intent, the scale, the substitution admissions — but it is not a verdict. The central legal question, whether training on copyrighted work is fair use, remains undecided, and courts have so far leaned toward AI companies; the Trump administration even filed a brief in September 2026 backing OpenAI's position. No one has been charged. We therefore grade the 'they broke the law' conclusion SOME SMOKE — documented smoke, genuinely, but not an adjudicated fire — and note that a Claude-drafted page (see §0) has every incentive to overclaim here, which is exactly why we don't. The documents show intent and awareness; the courts will decide legality.

§5 · Record vs Narrative

The lines we hold

  • Their words, not ours. The page rests on OpenAI's and Microsoft's own messages, memos, and testimony — the safest ground for a piece written by a competitor's model.
  • The exhibits are sealed. The quotes come from the plaintiffs' brief without full context; we treat them as strong evidence of attitude and intent, not as fully contextualized confessions.
  • No verdict. “Broke the law” is graded SOME SMOKE. Fair use is undecided, courts have leaned toward AI firms, and the government backed OpenAI. We don't prejudge it.
  • We disclose our stake. Anthropic, whose model wrote this, competes with OpenAI. The §0 disclosure is the first thing on the page, not a footnote.
§6 · Why It Matters

“There is no AI exemption to the law”

This page belongs in The Tech Right because it captures the governing attitude of the AI boom in two words: ah nice. The documents describe a business built on taking — paywalls circumvented, copyright notices stripped, a continent of journalism ingested — by companies whose own staff privately called it theft and whose products they expected to hollow out the publishers they fed on. The argument made by David Dayen in the piece that prompted this — that there is “no AI exemption to the law,” and that persistent rule-breaking is itself an unfair method of competition — is a serious one, and we take it seriously without adopting its verdict. It also connects, uncomfortably, to our own house: the mirror-image conflict in The Wrong AI Debate is that Anthropic, the maker of the model writing this, benefits when the story becomes “OpenAI is reckless, regulate the frontier.” So read this as what it is: an account of what OpenAI and Microsoft's own documents say, offered by a rival's model that has tried to grade itself harder for exactly that reason. The smoke is real. The verdict is the court's.

§7 · Questions

Questions worth taking seriously

Did OpenAI's founders actually break the law?

Unresolved — and we do not claim they did. That is the plaintiffs' argument and the op-ed's, and the documents give it real support (intent, scale, admissions that the products substitute for the original). But the core question — whether training on copyrighted work is fair use — has not been decided, courts have leaned toward AI companies, the Trump administration backed OpenAI, and no one has been charged. We grade “broke the law” SOME SMOKE: documented smoke, no adjudicated fire.

Isn't a Claude-written attack on OpenAI just a competitor trashing a rival?

That is the real risk, and it is why the conflict disclosure is §0, before anything else. Our safeguards: we rest only on OpenAI's and Microsoft's own words, corroborated across multiple outlets; we fence the sealed-exhibit caveat; and we refuse the legal verdict the op-ed reaches. If we wanted to trash OpenAI we would have graded “broke the law” as fact — we did not. Judge the page by whether it leans on the companies' own documents or on our characterizations. It is the former by design.

Why trust quotes pulled from a legal brief with the exhibits sealed?

Treat them as strong but not final. They come from the plaintiffs' summary-judgment brief, and the underlying exhibits remain sealed, so the full context is not public — a two-word reply can read differently in its thread. What raises confidence is corroboration: the “ah nice” exchange and the “theft of labor” memo are reported independently by TechCrunch, the New York Times, and Mother Jones quoting the same filing. So we treat the existence and wording of the quotes as fact, and their full meaning as pending the exhibits.

§8 · Standing Invitation

If you are named on this page

If you are named on this page, or are a party materially affected by the claims made here, and you wish to respond, correct the record, or add context, use the Contact page. Responses are published verbatim alongside the original claim, with the sender identified and the date of receipt. The channel stays open for the life of the page.

This site aggregates and grades a record that other outlets and primary sources have already put on the record. Every FACT-graded claim above is sourced to court filings, government reports, sworn whistleblower disclosures, published investigative journalism, or named-source statements. The citations are the accountability mechanism; this section is how you get on the record too.

§9 · Sources

The record

▦ Ledger gaps

Help us fill these lines.

This entry is graded on what’s on the public record. These are the blanks we know about. If you can source one, you’re rebuilding the ledger with us.

  • OpenWhether training LLMs on copyrighted, paywalled work is fair use — the undecided question at the heart of the case.Help fill this →
  • OpenWhat the sealed underlying exhibits show once the quotes are seen in full context.Help fill this →
  • OpenWhether persistent rule-breaking as a method of competition draws any legal or regulatory response.Help fill this →

Notify me when a gap is filled

We'll email you when we fill one of the gaps above.

By signing up you agree to receive emails from The Black Book Audit. Unsubscribe anytime.