Anthropic pays $1.5 billion for misusing books for AI training

Applications of AI


Reuters/Yonhap News - Seoul Economic News International news from South Korea
Reuters/Yonhap News

Anthropic has agreed to pay $1.5 billion (approximately 2.2 trillion won) after being embroiled in a copyright lawsuit for having its artificial intelligence (AI) model, Claude, study millions of books without permission.

According to Reuters, a US federal court in San Francisco on Thursday gave final approval to a $1.5 billion settlement between Anthropic and copyright owners. This is the largest known settlement in a U.S. copyright case and the first to reach an agreement in a lawsuit brought by copyright holders, including authors and news organizations, against an AI company over large-scale language model (LLM) training. A federal court tentatively approved the settlement last September.

In 2024, copyright holders filed a class action lawsuit alleging that Anthropic used the work in Claude’s training without permission.

Judge William Alsup, who was presiding over the case at the time, ruled last June that using the copyrighted work to train Claude constituted “fair use.” U.S. copyright law allows limited unauthorized use of copyrighted works for educational or non-commercial purposes, when the work is based on fact rather than fiction, and when the use does not impair the copyright owner’s potential revenue.

However, Judge Alsup found that Anthropic could be infringing because it kept digital copies of the books without legally purchasing them. He took issue with the company’s copying of more than 7 million books (not necessarily used for AI training) and storing them in a separate central digital library.

The case was originally scheduled to go to trial in December last year to determine the amount of damages. Some suggested damages could reach hundreds of billions of dollars. But the settlement reached between Anthropic and the plaintiffs gives rights to copyright holders in an unprecedented situation where copyrights were used for LLM training, while Anthropic, which plans to go public later this year, resolves potential risks, observers said.

Because this was an unprecedented copyright lawsuit, the number of works covered was enormous. Approximately 480,000 works were subject to court settlements, with more than 92% of copyright holders seeking compensation, according to plaintiffs’ lawyers. It is estimated that compensation of $3,000 (4.4 million won) will be paid for each book. However, some authors objected, arguing that the settlement amount was insufficient, that the plaintiffs’ attorney fees were excessive, and that some copyright holders were unfairly excluded from the settlement.

However, the court rejected all these objections in its ruling on the same day. It also approved payment of more than $101 million of the $187.5 million in fees charged by plaintiffs’ attorneys. Meanwhile, some authors and publishers who chose not to participate in the settlement continue to pursue separate lawsuits against Anthropic, Reuters reported.

Anthropic’s “Project Panama”…Cut out the spine and scan the pages.

Dario Amodei, Anthropic CEO. AFP/Yonhap News - Seoul Economic News International news from South Korea
Dario Amodei, Anthropic CEO. AFP/Yonhap News

Anthropic’s physical book training project is known as “Project Panama.”

The Washington Post reported that more than 4,000 pages of documents Anthropic filed with the court during the lawsuit included details of how the company used numerous books to train people. “Project Panama is our initiative to destructively scan every book in the world,” Anthropic said in a 2024 internal document, adding, “We do not want anyone to know we are running this project.”

Anthropic has consistently expressed a desire for quality writing that could improve Claude’s quality. In January 2023, WP reported that one of Anthropic’s co-founders claimed that training an AI model on books could teach it “actual writing,” rather than copying “low-quality internet terminology.” However, the company reportedly concluded that it would be virtually impossible to seek permission directly from copyright holders, and began acquiring books in bulk without their consent.

Anthropic hired Tom Turvey, who contributed to the Google Books project, and after considering other options, including approaching large used bookstores in New York and the New York Public Library, the company ended up purchasing millions of books. According to a report from WP, Anthropic was looking for a company that could scan between 500,000 and 2 million books in six months, using a cutting machine to cut the books and then scan them, according to a proposal from Anthropic’s partner companies.



Source link