OpenAI’s new media manager tool for creators

Machine Learning


OpenAI has sparked a heated debate over data privacy ever since ChatGPT was first made publicly available. The company used vast amounts of data from the public internet to train the large-scale language models that power ChatGPT and other AI products. However, this appears to have contained copyrighted content. Some creators have sued OpenAI, and several governments have launched investigations.

Basic privacy protections, such as opting out of training AI using your data, were missing even for regular users. Under pressure from regulators, OpenAI added a privacy setting that allows you to remove content to prevent it from being used to train ChatGPT.

Going forward, OpenAI will be introducing a new tool called Media Manager that will allow creators to opt out of training ChatGPT and other models that power OpenAI products. This feature may have been introduced much later than some expected, but it's still a useful privacy upgrade.

OpenAI published a blog post on Tuesday detailing its new privacy tool and how it trains ChatGPT and other AI products. Media Manager allows creators to identify content and tell OpenAI that they want it excluded from machine learning research and training.

Now, the bad news. This tool is not yet available. It is expected to be completed by 2025, and OpenAI plans to introduce additional options and features as development continues. The company also hopes to create new industry standards.

Still image from a video produced by OpenAI's Sora
Sora is OpenAI's AI-based text-to-video generator. Image source: OpenAI

OpenAI did not explain in detail how the Media Manager works. But it has big ambitions, as ChatGPT covers all types of content, not just text, that you might come across on the internet.

This includes cutting-edge tools to build a first-of-its-kind tool that helps identify copyrighted text, images, audio, and video across multiple sources and reflects the creator's preferences. Machine learning research will be required.

OpenAI also noted that it is working with creators, content owners, and regulators to develop media manager tools.

How OpenAI trains ChatGPT and other models

This new blog post didn't just announce a new Media Manager tool that could potentially prevent ChatGPT and other AI products from accessing copyrighted content. This can also be read as a declaration of the company's goodwill towards developing AI products that benefit users. And this sounds like a public defense against claims that ChatGPT and other OpenAI products may have used copyrighted content without permission.

In fact, OpenAI explains how to train the model and the steps to take to prevent unauthorized content and user data from entering ChatGPT.

The company also said it does not retain any of the data used to train its models. Models do not store data like databases. Also, each new generation of the underlying model obtains a new dataset for training.

Once the training process is complete, the AI ​​model does not retain access to the data analyzed during training. ChatGPT is like a teacher who can explain things because he has learned from many previous studies and the relationships between concepts, but does not store the content in his head.

OpenAI DevDay Keynote: How to use ChatGPT in 2023.
OpenAI DevDay Keynote: How to use ChatGPT in 2023. Image source: YouTube

Additionally, OpenAI stated that ChatGPT and other models should not regurgitate content. If this happens, it must be a mistake at the training level.

In rare cases, if the model incorrectly repeats representational content, it is a failure of the machine learning process. This failure is more likely to occur with content that appears frequently in the training dataset, such as content that appears on various public websites because it is frequently cited. We employ state-of-the-art technology to prevent repetition during training and when outputting the API or ChatGPT, and we continually improve it through continuous research and development.

The company also wants enough variety to train ChatGPT and other AI models. This means content in many languages, covering different cultures, subjects, and industries.

“Unlike the big companies in the AI ​​space, we don't have large amounts of data collected over decades. We rely primarily on publicly available information to teach our models how to be useful.” added OpenAI.

The company uses data “collected primarily from industry-standard machine learning datasets and web crawls, similar to search engines.” We exclude sources with paywalls, sources that aggregate personally identifiable information, and content that violates our policies.

OpenAI also uses data partnerships for content that is not publicly available, such as archives and metadata.

Our partners range from a leading private video library of images and videos training Sora to the Icelandic government helping to preserve the native language. We do not pursue paid partnerships purely for public information purposes.

The mention of Sora is interesting, as OpenAI has recently come under fire for failing to adequately explain how it trained the AI ​​models used in its sophisticated text-to-video products.

Finally, human feedback also helps train ChatGPT.

Regular ChatGPT users can also protect their data

OpenAI also reminds ChatGPT users that they can opt out of chatbot training. These privacy features already exist and predate the Media Manager tools currently in development. “Data from the ChatGPT Team, ChatGPT Enterprise, or our API platform” is not used for ChatGPT training.

Similarly, ChatGPT Free and ChatGPT Plus users can opt out of AI training.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *