TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Authors Guild says court filings in its lawsuit against OpenAI and Microsoft document internal discussions about training data from LibGen, a library of pirated books. The claims come from the authors’ legal brief; the allegations have not been established as court findings, and the case remains ongoing.
Authors suing OpenAI and Microsoft say newly filed court papers show that company employees discussed using books from LibGen, a source the papers describe as sketchy, and worried about how that use would look publicly. The claims put the companies’ knowledge and conduct at the center of a copyright case over books allegedly used to train AI systems; they are the plaintiffs’ account, not findings by a judge.
The Authors Guild said the materials were filed last week in Authors Guild v. OpenAI, part of multidistrict litigation pending in Manhattan. The cited documents include a motion for partial summary judgment and a corrected statement of undisputed material facts. The Guild’s September 21 release summarizes the plaintiffs’ arguments and quotes from the filings. OpenAI and Microsoft’s responses to these specific claims are not included in the material provided.
According to the plaintiffs’ brief, then-OpenAI Research Director Dario Amodei described LibGen as “a bit sketchier” as a training set. OpenAI researcher Sam McCandlish reportedly said he was worried about “optics,” including the prospect of a Hacker News post saying OpenAI used copyrighted data from a sketchy Russian website. The quoted exchange concerns employees’ reported discussion of public perception; by itself, it does not establish what legal conclusions they reached.
The filing also alleges that Microsoft knew OpenAI used LibGen by April 2019, when Sam Altman and Amodei presented an early GPT-3 version to Bill Gates and Microsoft technology chief Kevin Scott. The Guild says the documents describe an OpenAI effort called Project Clear and cite a June 2022 Slack exchange about removing LibGen references or files from company systems and storage. The materials supplied do not establish the full scope of any deletion or what data was used to train particular models.
The Dispute Over Training Data
The allegations matter because the lawsuit turns in part on whether copyrighted books were used to develop commercial AI systems, and what the companies knew about the source and status of that material. If substantiated, the communications could inform arguments over knowledge, decision-making and responsibility. Their legal significance will depend on the complete record and the court’s analysis; the excerpts alone do not decide whether training on the books infringed copyright.
The case also reaches beyond the parties. Authors and publishers have raised concerns that AI systems can produce text that competes with human writing, while AI developers have faced questions about the sources used to build training datasets. The Guild’s account links those broader concerns to internal comments about labor impacts. It quotes OpenAI policy director Jack Clark in 2020 warning that AI work could substitute for people’s labor and make genre-fiction authors worry about competition. Those remarks are presented by the plaintiffs as evidence of awareness, not as a court conclusion that particular writers lost work because of a specific product.
For readers, the distinction between a filing and a ruling is material. Court papers show what one side is asking the judge to accept and the evidence it cites. A partial-summary-judgment motion seeks resolution of specified issues; the supplied release does not state that the court has granted it.
How the Case Reached This Point
The Authors Guild and a group of writers brought claims on behalf of book copyright owners against OpenAI and Microsoft. Named authors include John Grisham, George R. R. Martin, Jodi Picoult and others. The Guild describes the action as part of broader multidistrict litigation against the companies in Manhattan. The plaintiffs allege that copyrighted books were used without permission in training AI products. Those allegations remain contested claims in litigation.
The Guild’s release says its latest brief accompanies separate filings from news media organizations and follows earlier revelations in the case. It frames the new material as evidence that company leaders and employees knew about possible harms to writers and made deliberate choices about data. That characterization belongs to the plaintiffs. The excerpts provided do not include the defendants’ full account of how training datasets were assembled, their legal arguments, or any judicial assessment of the evidence.
One portion of the plaintiffs’ account concerns Tarun Gogineni, whom OpenAI hired in 2022 to improve its models’ writing quality. The filing quotes him describing a research ambition for GPT models to complete the final books of Martin’s A Song of Ice and Fire series. It also attributes to him remarks about authors’ complaints and economic disruption. The Guild uses these comments to argue that employees recognized the prospect of competition with writers. The quoted statements do not establish that OpenAI released a product that completed Martin’s series or that the company adopted the employee’s stated view as policy.
“These filings reveal shocking disdain for writers and their work through repeated, intentional decisions to steal books rather than pay for them.”
— Authors Guild CEO Mary Rasenberger
Questions the Filings Leave Open
The supplied source is the Authors Guild’s account of filings by the plaintiffs. It does not include responses from OpenAI or Microsoft, the complete exhibits, or a court ruling on the claims described. The defendants’ positions and any disputes over the meaning or completeness of the quoted communications are therefore not clear from this material.
It is also unclear from the release which specific books were in LibGen materials used for training, which models or training runs involved them, and what effect the reported discussions had on company decisions. The Guild says OpenAI deleted LibGen files in 2022 under Project Clear, but the information provided does not establish the exact files removed, whether copies remained elsewhere, or whether deletion affected previously trained models. The precise questions before the judge on partial summary judgment are not detailed in the release.
Briefing and Hearing Ahead
The Authors Guild expects additional briefing over the next several months and a hearing in early 2027, according to its September 21 announcement. The court will consider the parties’ filings and evidence before deciding any issues presented for summary judgment. The announcement does not provide a hearing date or predict when the judge might rule.
Further filings may clarify how OpenAI and Microsoft respond to the allegations, what evidence they dispute, and how the plaintiffs connect internal communications to the legal questions in the case. Until the court rules, the claims about LibGen use, employee knowledge and file removal remain allegations advanced by the authors’ side.
Key Questions
What did the authors’ filing say about Hacker News?
The plaintiffs’ filing quotes OpenAI researcher Sam McCandlish saying he worried about the “optics” of a public account that OpenAI used copyrighted data from a sketchy Russian website. The quote records a reported concern about perception; it is not a judicial finding.
What is LibGen?
LibGen is the name used in the filings for a digital library that the Authors Guild describes as a source of pirated books. The release says OpenAI employees discussed using material from it as training data.
Has a court found that OpenAI or Microsoft broke copyright law?
The source material describes allegations in plaintiffs’ court papers, not a ruling that either company infringed copyright. The case is ongoing.
What did the filing allege about Microsoft’s knowledge?
According to the plaintiffs, Microsoft knew about OpenAI’s use of LibGen by April 2019, when an early GPT-3 presentation was made to Bill Gates and Microsoft technology chief Kevin Scott. The supplied release does not include Microsoft’s response.
When could the court consider the claims?
The Authors Guild says it expects more briefing over the next several months and a hearing in early 2027. It has not provided a specific hearing date.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
