This was a highly controversial ruling that could have changed the course of this and other lawsuits against AI companies. The court had just ruled that OpenAI waived attorney-client privilege, rejecting claims that it knowingly infringed the copyrights of authors of illegally downloaded books. The discovery opens the door to internal communications behind the company’s deletion of two huge data sets of pirated books, potentially exposing the company to huge losses. The company’s in-house legal team was scheduled to be fired.
OpenAI immediately appealed the ruling, suing Lisa Blatt, a veteran Supreme Court lawyer whose clients include Google, Bank of America, and Starbucks. In her brief, she issued a dire warning. If left in place, the decision would eviscerate privilege claims in all copyright cases involving so-called mental states, an analysis used to determine whether a defendant intentionally infringed or was unaware that he or she was infringing.
On Friday, OpenAI won its argument to overturn the court’s ruling. Billions of dollars were at stake. This communication can help prove willful infringement, which could result in damages of as much as $150,000 instead of just $200 per work. And perhaps more importantly, the decision threatens to provide companies suing AI companies with an avenue to obtain evidence that would normally be considered privileged information.
This issue has become an important battleground in discovery. This relates to an OpenAI employee downloading a pirated copy of the book in 2018 and using it to create two datasets known as “Book 1” and “Book 2” to train two obsolete GPT models. The company initially told the court that the data sets were deleted in 2022 “because they were not used,” but has since maintained that information about the reason for the deletion was kept secret. Lawyers representing the author and publisher alleged fraud.
In November, Magistrate Judge Ona Wang ruled that OpenAI must provide evidence to clarify its motives for deleting its datasets. She concluded that the company opened the door to privileged material when it revealed that “Book 1” and “Book 2” were removed for “non-use,” another part of her order that caught the attention of copyright courts. Mr. Wang reasoned that the company had effectively waived attorney-client privilege by denying the willful infringement claim, which he said would bring the company’s mental state into the spotlight in court. “Denying that OpenAI knowingly infringed Class Plaintiffs’ copyrighted works amounts to asserting that OpenAI acted in good faith,” according to the order.
In Friday’s order, which ruled against authors and publishers including Sarah Silverman, U.S. District Judge Sidney Stein emphasized that denying willful infringement claims is not the same as advancing a good faith defense. This would take finding out why OpenAI wiped its dataset in 2022 squarely off the table, she said. “There is a difference between a copyright defendant who simply denies a claim of intent, where the burden of proof rests on the plaintiff, and a copyright defendant who affirmatively asserts a good faith belief that his actions are lawful,” Stein wrote.
Due to the reversal, the discovery battle becomes a preliminary detour of the incident. The ruling was overturned, but the argument was a sly one by lawyers representing the authors, led by Justin Nelson and Craig Smither of Sussman Godfrey, the firm that negotiated the $1.5 billion Anthropic settlement. If the ruling stands, AI companies would be burdened with the burden of proving they did not intend to violate copyright law every time they deny a claim that they knowingly infringed a copyrighted work.
The revocation also discussed whether OpenAI disclosed privileged information when it said the dataset was deleted “because it was not used.” On this issue, Stein said that this assertion does not represent legal advice and cannot be used as a basis for finding that OpenAI has waived privilege.
Despite losing in the battleground of discovery, the authors’ lawyers have gained the upper hand in a debate over the piracy of books from shadow libraries that is increasingly showing signs of victory. This theory has changed over the course of AI litigation. Initially, the authors’ lawyers directly tied the copyright infringement to OpenAI’s training of models under a single umbrella. But the two men later separated their theories, arguing that the clear act of illegally downloading a work constitutes copyright infringement, regardless of whether it is used or not.
The move capitalizes on one victory for the author in another AI copyright case. The lawsuit was filed by Andrea Bartz against Anthropic and involves the company illegally downloading and storing millions of books in its central library. Although the ruling was heavily tilted in Anthropic’s favor, the court gave the green light to consider the theory, which is now part of the OpenAI lawsuit. “Anthropic’s subsequent purchase of copies of books it previously stole from the Internet does not relieve it of liability for the theft,” said U.S. District Judge William Alsup. After the verdict, Anthropic agreed to pay $1.5 billion to settle the case.
What today’s AI systems are trained on remains largely unknown. OpenAI trained an older version of GPT using “Book 1” and “Book 2” downloaded from the Shadow Library website LibGen, but later deleted the dataset. Still, AI companies have raised billions of dollars largely thanks to models trained on pirated books.
