Chatbot Builder Humanity has agreed to pay the author $1.5 billion in a groundbreaking copyright settlement that will allow artificial intelligence companies to redefine the way they compensate creators.
The San Francisco-based startup is ready to pay authors and publishers, and is settling lawsuits accusing the company of using its work to train chatbots.
Humanity has developed an AI assistant named Claude that can generate text, images, code, and more. Writers, artists and other creative experts raised concerns that humanity and other tech companies use their jobs to train AI systems without permission and do not compensate them fairly.
As part of the settlement that the judge still needs to approve, humanity has agreed to pay the author $3,000 per job for an estimated 500,000 books. This is the biggest settlement known to signal copyright to other high-tech companies facing allegations of copyright infringement.
ChatGpt manufacturers Meta and Openai are also sued for alleged copyright infringement. Walt Disney Co. and Universal Pictures are suing AI company Midjourney. This claims that the studio trained image generation models with copyrighted material.
“It provides meaningful compensation for each class of work and sets precedents that require AI companies to pay the copyright owner,” author's lawyer Justin Nelson said in a statement. “This settlement sends a strong message to AI companies and creators at the same time that it's wrong to take copyrighted works from these pirate websites.”
Last year, authors Andrea Burtz, Charles Graeber and Kirk Wallace Johnson sued humanity, claiming that the company committed “massive theft” and trained chatbots with pirated copies of copyrighted books.
San Francisco US District Judge William Alsp ruled in June that it was not illegal because humanity used books to train AI models to form “fair use.” However, the judge also ruled that the startup had improperly downloaded millions of books through online libraries.
Fair use is a legal tenet in US copyright law, allowing limited use of copyright devices in certain cases, such as education, criticism, and news reporting. When AI companies are sued for alleged copyright violations, they point out their doctrine as a defense.
Founded by former Openai employees and backed by Amazon, humanity pirated at least 7 million books from online libraries, including Books3, Library Genesis and Pirate Library Mirror, and online libraries containing fraudulent copies of copyrighted books, according to the judge.
He also bought millions of print copies in bulk, stripped off the bindings from the book, cut pages and scanned them into digital and machine-readable forms.
In subsequent order, Alsup pointed to the potential damages of copyright holders of books downloaded by mankind from Shadow Libraries Libgen and Pilimi.
The award was massive and unprecedented, but according to some calculations it could have been much worse. If humanity was charged the largest penalty for each of the millions of works it used to train AI, the bill could have been more than $1 trillion, some calculations suggest.
Humanity opposed the ruling and did not admit fraud.
“Today's settlement will resolve the remaining claims of the plaintiffs, if approved,” Aparna Sridhar, deputy adviser of humanity, said in a statement. “We are continuing to be committed to developing secure AI systems that help people and organizations expand their capabilities, advance scientific discoveries and solve complex problems.”
The human conflict with the author is one of many cases where artists and other content creators are challenging the companies behind AI generated to compensate for the use of online content to train AI systems.
The training will provide a vast amount of data, including social media posts, photos, music, computer code, videos, and more, to train AI bots to identify languages, images, sounds, and conversation patterns that can be mimicked.
Some tech companies have won copyright lawsuits filed against them.
In June, the judge dismissed the author of a lawsuit filed against Facebook's parent company Meta. This claimed that they also developed AI assistants and stole work to train AI systems. US District Judge Vince Chhabria said the lawsuit was thrown because the plaintiffs “had a false argument,” but the ruling “did not support the proposal that it was legal for Meta to train language models using copyrighted material.”
The trade group, which represents the publisher, praised the human settlement on Friday, noting that it would send a big signal to high-tech companies developing powerful artificial intelligence tools.
“Beyond financial terms, the proposed settlement offers great value when sending messages that artificial intelligence companies cannot illegally obtain content from Shadow Library or other pirate sources as building blocks for models.”
The Associated Press contributed to this report.
