Microsoft AI CEO Mustafa Suleiman said this week that machine learning companies can scrape most of the content publicly available online and use it to train neural networks because it's essentially “freeware.”
Shortly thereafter, the Center for Investigative Journalism sued OpenAI and its largest investor, Microsoft, for using content from nonprofit news organizations without permission and offering compensation to them.
This follows a lawsuit filed in April by eight newspapers against OpenAI and Microsoft for misappropriating their content, four months after The New York Times also filed a lawsuit.
Additionally, two authors sued OpenAI and Microsoft in January, alleging that the authors' copyrighted work was used to train AI models without their permission, and in 2022, several unidentified developers sued OpenAI and GitHub, alleging that the companies used publicly available programming code to train generative models in violation of software license terms.
Asked in an interview with CNBC's Andrew Ross Sorkin at the Aspen Ideas Festival whether AI companies are effectively stealing the world's intellectual property, Suleiman acknowledged the debate but sought to distinguish between content that people post online and content that is backed by corporate copyright holders.
“When it comes to content that's already on the open web, I think the social contract around that content has been fair use since the 1990s,” he said. “Anyone can copy it, recreate it, reproduce it. It was freeware, so to speak. That's the understanding.”
Suleiman acknowledged that there is another category of content: content published by companies staffed by lawyers.
“There's another category where websites or publishers or news organizations can specifically say, 'Please don't scrape or crawl for any purpose other than indexing,' so that other people can find that content,” he explained. “But it's a gray area. And I think that will be resolved in court.”
That's an understatement. While Suleiman's comments are sure to anger content creators, he's not entirely wrong. It's not clear where the legal lines are when it comes to training AI models and the output of those models.
Most people who post content online as individuals are violating rights in some way by agreeing to the terms of use offered by the major social media platforms. Reddit's decision to license its users' posts to OpenAI would not have been made if the company believed its users had legitimate rights to their memes and manifestos.
The fact that OpenAI and other AI modeling companies have struck content deals with major publishers shows that a strong brand, deep pockets and legal team can bring big tech businesses to the negotiating table.
In other words, anyone who creates content and posts it online is creating freeware, unless they can hire or attract lawyers willing to challenge Microsoft and companies like it.
In a paper published on SSRN last month, Frank Pasquale, professor of law at Cornell Tech and Cornell Law School in the US, and Hao-Cheng Sun, associate professor of law at the University of Hong Kong, explore the legal uncertainty surrounding the use of copyrighted data to train AI and whether courts would find such use fair. They conclude that AI needs to be addressed at the policy level, as current law is not suited to answer the questions that need to be addressed now.
“Given significant uncertainty about the lawfulness of AI providers' use of copyrighted works, lawmakers need to articulate a bold new vision for rebalancing rights and responsibilities, just as they did with the development of the internet (the Digital Millennium Copyright Act of 1998),” they argue.
The authors suggest that continued free collection of creative works will threaten not only writers, composers, journalists, actors and other creative professionals, but also generative AI itself, eventually leading to a shortage of training data: they predict that people will stop putting their work online if it is used to power AI models that reduce the marginal cost of content creation to zero, removing the possibility of rewarding creators.
That's the future Suleiman sees. “The economics of information are going to change fundamentally because we can bring the cost of producing knowledge down to zero marginal cost,” he said.
All of this freeware that you helped create is available for a small monthly subscription fee.®
