Newly unsealed court material in a copyright lawsuit filed in 2023 has put unusually blunt internal discussions from Microsoft and OpenAI personnel into public view. The documents concern the acquisition and use of news material for training large language models, and the excerpts released so far depict employees and executives wrestling with a central contradiction of generative AI: a system may be built from a vast supply of human-made work while also making that work less valuable to the people and organisations that create it.
One Microsoft applied-science director, Dr. Brent Hecht, reportedly described the practice in exceptionally severe terms, calling it the “largest theft of labor in human history” and warning of a potential “doom loop.” OpenAI executive Nick Turley reportedly described the prospect as an “existential threat to publishers.” Those remarks are significant because they suggest that concerns about the effect on journalism were being voiced inside companies involved in building and deploying the technology, not solely by publishers challenging it in court.
Important caution is warranted, however. Only snippets of the newly unsealed material are public, without the surrounding conversations or complete context. The excerpts provide evidence of particular views expressed internally; they do not, by themselves, settle what either company’s overall policy was, whether every employee agreed, or whether the conduct at issue violated copyright law. Those are precisely the kinds of questions the litigation is meant to test.
What the lawsuit is about
The case was brought by a major newspaper publisher alongside five writers. Their core allegation is that AI companies took millions of articles from across the web without permission or payment and used that writing to train advanced language models. The dispute is not a narrow argument about a single copied story. It concerns the input layer of modern generative AI: the enormous bodies of text used to help systems predict and generate language.
A large language model, or LLM, is a system trained on text to identify patterns in language and produce responses. It does not function like a searchable shelf of articles in the ordinary sense. But the legal and economic dispute is not erased by that distinction. Copyright holders are asking whether copying their work into training datasets is lawful in the first place, and whether tools that answer users’ questions can become substitutes for the original reporting.
The plaintiff argues that the companies accessed stories at industrial scale, including by bypassing paywalls, and that copyright notices were removed from training material. The newly available documents reportedly provide detail on the methods used by OpenAI and its partners to obtain content, including material behind paywalls and datasets composed of millions of documents.
That matters because paywalls are not merely a browser annoyance. They are part of the mechanism through which publishers finance reporters, editors, fact-checking, travel, legal review and the many less visible costs of newsgathering. Circumventing one, if established, is a materially different allegation from reading text that a publisher has made freely accessible to everyone.
The “doom loop” explained
Hecht’s reported “doom loop” warning captures the practical fear behind the lawsuit. Journalism requires investment to create: people investigate, interview sources, corroborate claims and publish the result. A generative AI service can then use that reporting as input, deliver a condensed answer to a user, and reduce the chance that the user visits the publisher’s page. If fewer visits mean less income, the publisher may be less able to fund the journalism that gave the AI system useful material in the first place.
It is a supply-chain concern, but not in the usual factory-and-shipping sense. The “supply” is original reporting. In another reported comment, Hecht said large AI models could be “a product that destroys its supply chain.” The phrase is memorable because it identifies the economic question more clearly than a fight over technical jargon: can a product thrive by consuming the work of an ecosystem it simultaneously weakens?
Turley’s reported assessment that AI products are “largely substitutive” to journalism points at the same issue. A substitute is something users choose instead of another product. An AI answer can be helpful, but if it satisfies the user’s need for a summary, explanation or current fact without sending meaningful attention back to the reporting, it may compete with the publisher rather than refer readers to it.
A reported 2023 message from an OpenAI engineer makes that concern even more direct: “no matter how prominently we show the links, users won’t click.” Links can be valuable for verification, attribution and discovery. But a link is not automatically a business model. A user who gets a complete-seeming answer on an AI interface may have little reason to open the cited page, particularly if they only wanted a quick response. The question is therefore not only whether a link appears, but whether the design leaves publishers with a meaningful role in the reader relationship.
Why the paywall allegations stand out
Among the reported details, the allegations involving paywalls are especially consequential. A message described an employee telling OpenAI president Greg Brockman about a new paywall hack, with Brockman replying, “ah nice.” The available excerpt is brief and lacks broader context, so it should not be expanded beyond what it says. Still, it is likely to attract close attention because it concerns intentional access controls rather than the broader and more contested question of whether web scraping of publicly reachable pages can qualify as fair use.
Similarly, the reported removal of copyright notices from training data could matter independently of the broader training debate. Copyright notices identify ownership and usage rights. The case will need to establish the specific facts and determine their legal significance, but the allegation helps explain why the litigation has become so important: it combines arguments about large-scale copying, access to restricted content, attribution information and downstream competition with news products.
Fair use is the legal hinge, not a blank cheque
The central legal doctrine in play is fair use. In broad terms, fair use can permit copyrighted work to be used without permission in certain circumstances, including examples such as commentary, parody and journalism. AI companies have argued that training is transformative: the text is used to develop a new system rather than republish an article as an article.
“Transformative” is a legal concept with major stakes here. It generally asks whether a new use has a different purpose or character from the original. But calling a use transformative does not automatically end a copyright analysis. The lawsuit raises difficult questions about the quantity of material taken, the nature of the copying, how the material was obtained, and—especially—whether AI outputs or AI search-like experiences damage the market for the originals.
Some AI-related cases have already produced rulings favourable to AI companies. Yet judges have also emphasised that the legal rules around AI uses are not fully settled. That makes this lawsuit a major test of whether developers can rely on fair use when they take publishers’ material for training and then provide tools that may compete for the same reader attention.
The legal outcome cannot be read from an inflammatory internal message. Courts assess evidence, claims and defences in full. But internal descriptions can be relevant to the public debate because they reveal how people inside a company framed the possible consequences of a strategy. The gap between private concern and public legal position is likely to be a focal point as the case continues.
Microsoft’s stated position
Microsoft has disavowed the views expressed by individual employees in the materials. Its stated position in the case is that the uses at issue are transformative and consistent with copyright law, and that Copilot is not a replacement for publishers’ journalism. That is a clear and important distinction: private messages attributed to staff are not the same as the company’s legal argument.
At the same time, the comments reported from inside Microsoft raise a broader governance question. If internal staff identify a risk that AI could reduce the economic incentive to produce reliable information, companies have to decide whether links and general assurances are enough—or whether licensing, compensation, access restrictions, product design and data-provenance rules need a larger role.
Data provenance means knowing where training material originated, what rights attach to it and how it was acquired. It is not a glamorous feature for end users, but it may become one of the most consequential operational issues in AI. A company that cannot clearly track its data faces more uncertainty when a creator or publisher challenges its use. A company that can track it may be better positioned to license material, exclude restricted work or explain how it handled rights information.
Why this reaches beyond publishers
The immediate dispute is about journalism, but its eventual reasoning could affect a far wider set of creators whose work exists online. The underlying tension—innovative tools drawing value from existing labour while potentially redirecting audiences away from it—is not confined to one medium. It also matters to game writers, artists, guide makers, developers and communities whose work can be indexed, summarized or replicated in part by automated systems.
For the games business, the policy conversation arrives alongside other major platform questions, from user trust to the cost of abuse prevention. For example, industry estimates around cheating show how technology companies can face incentives and harms that extend well beyond a single product feature, as covered in this look at the broader video-game cheat economy. Copyright and AI are a different problem, but both require companies to account for effects that may be distributed across creators, players and platforms.
For readers, the most practical takeaway is to distinguish three separate issues that are often bundled together: whether an AI system can technically summarize information, whether it was legally trained on the material that informed it, and whether its presentation sends value back to original creators. A useful answer with citations can still raise questions about training data. A lawful training practice, if a court finds one, would not automatically prove that every product design treats publishers fairly. And a publisher’s understandable concern about lost traffic does not itself decide the law.
The unsealed excerpts sharpen those distinctions rather than resolve them. They show reported internal recognition of the possibility that generative AI could undercut journalism’s business model, while Microsoft maintains that its products are lawful transformative tools rather than substitutes for reporting. The court case will determine far more than the meaning of a few messages. It may help define what obligations, if any, AI developers owe to the human-made information systems on which their models depend.






