Meet C-suite in San Francisco July 11-12 to learn how business leaders are ahead of the generative AI revolution. learn more
Lawsuits targeting data-scraping practices by AI companies developing large-scale language models continue to rage today, with comedian and author Sarah Silverman writing in her humorous memoir, The Bedwetter: News has surfaced that they are suing OpenAI and Meta for copyright infringement of Stories of Courage, Redemption. and Pea, published in 2010.
The lawsuit, filed by San Francisco-based Joseph Saberi law firm, which also filed a lawsuit against GitHub in 2022, said Silverman and two other plaintiffs allege copyright It claims that it did not agree to use the protected books as training materials for OpenAI’s ChatGPT and Meta. When LLaMA and ChatGPT or LLaMA are requested, the tool will generate a synopsis of the copyrighted work. This is only possible if the model was trained on them.
>>Follow VentureBeat’s Coverage of Continuously Generated AI<
Legal AI issues over copyright and ‘fair use’ grow
These legal issues around copyright and “fair use” aren’t going away. In fact, they are the building blocks of today’s Large Language Models (LLMs): the heart of the training data. As we discussed last week, web scraping of large amounts of data is perhaps the secret sauce of generative AI. AI chatbots such as ChatGPT, LLaMA, Claude (Anthropic), and Bard (Google) are trained on large data corpuses collected primarily from the internet, so they can spit out consistent text . And as the size of his LLMs today, such as GPT-4, swells to hundreds of billions of tokens, so does the hunger for data.
event
transform 2023
Join us July 11-12 in San Francisco. There, he shares how management integrated and optimized his AI investments to drive success and avoid common pitfalls.
Register now
Data scraping practices under the guise of AI training have recently come under attack. For example, OpenAI was hit by two new lawsuits from his other. His one case, also filed on June 28 by Joseph Saveri Law Firm, found that OpenAI did not obtain consent from the copyright holders, did not provide credit or compensation, and filed texts in the book. alleged to have been illegally copied. Another, filed the same day by Clarkson Law Firm on behalf of a dozen anonymous plaintiffs, shows OpenAI’s ChatGPT and DALL-E violating privacy laws to collect people’s personal data from the internet. claims to be.
The lawsuits follow a class action lawsuit filed in January. Andersen et al. v. Stability AI, The plaintiffs in the lawsuit allege copyright infringement, and Getty Images filed a lawsuit against Stability AI in February, alleging copyright and trademark infringement and trademark dilution.
Sarah Silverman, of course, adds another celebrity layer to the AI and copyright debate, but what does this new lawsuit really mean for AI? is.
1. More lawsuits are coming.
Margaret Mitchell, researcher and chief ethics scientist at Hugging Face, wrote in my article last week that she called the AI data scraping problem a “pendulum swing” that would force OpenAI to take it down by the end of the year. He added that he had previously predicted that there might be These data issues require at least one model.
Certainly, we should expect to see more lawsuits. Back in April 2022, when the DALL-E 2 first appeared, Marc Davis, a partner at San Francisco-based law firm Orrick, said there was an open debate when it came to AI and “fair use.” Agreed that there are many legal issues. Promotes freedom of expression by permitting unauthorized use of copyrighted works in certain circumstances. “What really happens is when there’s a lot of stake involved, you file a lawsuit,” he said. “And you get an answer for each case.”
And now, a new debate around data scraping is “percolating,” Gregory Leighton, a privacy law expert at law firm Porcinelli, told me last week. He said the OpenAI lawsuit alone would be such a flashpoint that further backlash was inevitable. “We are less than a year old in the large language model era.
Bradford Neumann, head of the machine learning and AI practice at global law firm Baker McKenzie, said last October that the legal battle over copyright and fair use could ultimately go to the Supreme Court. told me there is.
“Legally, at this time, there is little guidance on whether input to copyrighted LLM training data is ‘fair use,'” he said. “I think it will end up in the Supreme Court,” he said, expecting different courts to come to different conclusions.
2. Datasets will come under increasing scrutiny, but it will be difficult to enforce.
In Silverman’s lawsuit, the authors allege that OpenAI and Meta deliberately removed copyright management information, such as copyright notices and titles. “Meta believes that this deletion will [copyright management information] By hiding the fact that all output from the LLaMA language model is an infringing derivative work, you promote copyright infringement,” the authors argued in their complaint against Meta.
The authors’ complaint also speculates that ChatGPT and LLaMA may have been trained on large datasets of books to circumvent copyright laws, including “shadow libraries” such as Library Genesis and ZLibrary. .
“These shadow libraries have long been of interest to the AI training community because they host a large amount of copyrighted material,” the authors’ complaint against Meta states. “Therefore, these shadow libraries are also grossly illegal.”
But a Bloomberg legal article last October pointed out that there are a number of legal hurdles that must be overcome when fighting shadow libraries for copyright. For example, according to Jonathan Band, intellectual property attorney and founder of Jonathan Band PLLC, many of the site operators are based outside the United States.
“They are beyond the scope of US copyright law,” he said in the article. “It is theoretically possible to go to the country where the database is hosted. It also raises various questions about whether we have a functioning justice system.”
Additionally, the burden is often on the author to prove that a “derivative” work was created as a result of using a copyrighted work for AI training. In an article for The Verge last November, Vanderbilt Law School professor Daniel Gervais argued that while it’s probably legal to use copyrighted data to train generative AI, the same isn’t necessarily true. said not. generate Content — that is, anything you do using that model can be infringing.
And Katie Gardner, a partner at international law firm Gunderson Dettmer, told me last week that fair use is “a defense against copyright infringement, not a legal right.” Moreover, it can also be very difficult to predict what courts will decide in a particular fair use case, she said, adding, “There is precedent for two cases with seemingly similar facts to rule differently. There are many,” he said.
But she argued that Supreme Court precedent, in which many reasoned that the use of copyrighted material to train an AI, could be fair use based on the transformative nature of such use I emphasized one thing: it’s not a transplant of the original work market.
3. Companies will want their own models and compensation.
Enterprises have already made it clear that they do not want to deal with the litigation risks associated with AI training data. You want secure access to create risk-free generative AI content for commercial use.
Compensation was the focus. Last week, Shutterstock announced that it would provide commercial customers with full indemnification for the licensing and use of AI-generated images on its platform to protect them from potential claims related to their use of the images. The company said it will respond to claims for compensation upon request with human review of the images.
The announcement comes just a month after Adobe announced a similar service. “If a customer is sued for infringement, Adobe will take over legal defense and provide some financial compensation for the claim,” said a company spokesperson.
And new poll data from enterprise MLOps platform Domino Data Lab shows that data scientists believe generative AI will have a major impact on enterprises in the next few years, but that function cannot be outsourced. It turns out that companies have to fine-tune or control generative AI on their own. AI model. In addition to data security, intellectual property protection is another issue, said Kjell Carlson, head of data science strategy at the Domino Data Lab. “If it’s important and really drives value, they want to own it and have a higher degree of control over it,” he said.
VentureBeat Mission will be the digital town square for technical decision makers to learn and transact on transformative enterprise technologies. Watch the briefing.
