See how companies are responsibly integrating AI into production environments. This invite-only event in SF explores the intersection of technology and business. Click here to learn how to participate.
OpenAI has been rolling out new updates at a rapid pace today alone, but the biggest one may be a new tool the company is developing called Media Manager. Scheduled to be released next year in 2025, creators will be able to select their own work if they have one. You will be able to scrape and train for the company's AI models.
Announced in a blog post on the OpenAI website, the tool is described as:
“OpenAI is developing Media Manager, which allows creators and content owners to tell us what they own and specify how their work is included or excluded from machine learning research and training. It's a tool that can. We plan to introduce additional options and features in the future.
This includes cutting-edge tools to build a first-of-its-kind tool that helps identify copyrighted text, images, audio, and video across multiple sources and reflects the creator's preferences. Machine learning research will be required.
VB event
AI Impact Tour – San Francisco
request an invitation
We work with creators, content owners, and regulators to develop Media Manager. Our goal is to have this tool in place by 2025, and we hope it becomes the standard across the AI industry.“
There is no price listed yet for this tool, but since OpenAI is using it to establish itself as an ethical actor, it will likely be provided for free.
The tool aims to provide creators with additional protection against AI data scraping beyond adding a code string to the robots.txt file on their website (“User agent: GPTBot Disallow: /”) . This is a measure that OpenAI introduced in his August 2023.
After all, many creators post their work on sites they don't own or control (platforms like DeviantArt or Pateron), where they can't edit the robots.txt file on the page. Additionally, some creators may want to exclude only certain works from AI data scraping and training, rather than all of their submissions, so OpenAI's proposed Media Manager makes this kind of more Allows for fine-grained control and options.
Additionally, OpenAI notes that in domains where opt-out text is not provided, creators' work may be easily screenshotted, stored, reshared, and otherwise reposted or redistributed on the web. It points out that there is.
“Many creators do not control the websites on which their content appears, and their content is often quoted, reviewed, remixed, reposted, and used as inspiration across multiple domains, making this difficult. We understand that this is an imperfect solution. We need an efficient and scalable solution that allows content owners to express their preferences for the use of their content in AI systems.”
Responding to strong and persistent criticism of AI data scraping
The move comes as AI model makers like OpenAI and its rivals Anthropic, Meta, and Cohere are scraping the web to obtain data for training without explicit permission, consent, or compensation. It comes amid a continuing wave of visual artists and creators protesting.
Several creators have filed a class action lawsuit against OpenAI and other AI companies, alleging that this data scraping practice infringes the copyrights of their images and works.
OpenAI's defense is that web crawling and scraping has been accepted and standard practice among many companies on the web for decades, an argument we revisit in today's blog post. Alluding to this, he writes: It was also voluntarily adopted by the Internet ecosystem to show web publishers which parts of his website can be accessed by web crawlers. ”
In fact, many artists tacitly accept scraping their data for indexing by search engines like Google, but oppose the resulting generative AI training. Because it competes more directly with their own work and livelihood.
OpenAI is offering indemnification (guarantee of legal aid and defense) to subscribers of its paid plans accused of copyright infringement, reassuring its growing list of lucrative corporate customers.
Ongoing legal issues
Courts have yet to rule definitively on whether AI companies and others can scrape copyrighted works without the creator's explicit consent or permission. But clearly, regardless of how the legalities play out, OpenAI wants to position itself as a collaborative and ethical organization when it comes to creators and their data sources.
That said, creators are likely to view this move as “too little, too late.” Because much of the creator's work has probably already been scraped and used to train AI models, and nowhere has OpenAI hinted that it may or will remove any of its models. Because there isn't. I was trained in such works.
In a blog post, OpenAI says that rather than storing copies of scraped data at scale, it only stores “equations that best describe the relationships between words and the underlying processes that generate them.” It is claimed that
The company writes:
“We design our AI models as learning machines, not databases.
Our model is designed to help you generate new content and ideas, rather than repeating or “regurgitating” content. AI models can describe facts that are in the public domain. In rare cases, if the model incorrectly repeats representational content, it is a failure of the machine learning process. This failure is more likely to occur with content that appears frequently in the training dataset, such as content that appears on various public websites because it is frequently cited. We employ state-of-the-art technology to prevent repetition during training and when outputting the API or ChatGPT, and we continually improve it through continuous research and development.“
At the very least, a media manager tool could be a more efficient and user-friendly way to block AI training than existing options such as Glaze or Nightshade. However, if it comes from OpenAI, it's not yet clear whether the creator can trust it. It also doesn't know if it can block training by other rival models.
