A major AI research lab and a major tech company are being accused of using subtitles from tens of thousands of YouTube videos without permission to train their artificial intelligence models.
Google has strict rules that prohibit it from scraping material from YouTube without permission. ProofNews Research The investigation found that companies using subtitles for over 170,000 videos include Apple, Nvidia, and Anthropic.
The captions are part of “The Pile,” a massive dataset compiled by nonprofit EleutherAI that was originally intended to give small businesses and individuals a way to quickly train models, but which has since been adopted by large tech and AI companies.
While Apple, Nvidia, and Anthropic did not scrape YouTube videos directly, the AI models the company operates, including Claude and Apple Intelligence, used Pile as a source of information and were therefore trained on that information.
Hunger for data

Several studies have found that two ingredients are essential to creating more advanced AI models: data and computing power.
Increasing either or both can improve responsiveness and improve performance and scale, but data is becoming an increasingly scarce and expensive commodity.
There are currently several lawsuits being filed against AI image and music generation companies over whether their training data constitutes fair copyright use.
Companies like OpenAI and Google are combining their own large data repositories with deals with major publishers and Reddit.
Meta has Facebook, Instagram, Threads and WhatsApp but has faced backlash from users. Apple has a huge amount of user data, but its own privacy policy makes it less useful for early model training.
The lack of available data has forced companies to look for new sources to train their next generation models, but not all of these sources are willing to provide data or are aware that the information they are producing is being used to train AI.
There are currently several lawsuits being filed against AI image and music generation companies over whether their training data constitutes fair copyright use.
What went wrong?

While Apple and Anthropic are not directly responsible for the use of these YouTube captions in model training datasets, their inclusion raises questions about the provenance of data and how strict the tech giants are when assessing rights.
It's not just videos from smaller creators that were included: videos from the BBC, NPR, The Wall Street Journal, MrBeast, and Marques Brownlee were all included in the dataset.
Nebula CEO Dave Wiskas said using the data without consent was “theft” and “disrespectful,” especially since the studio is already aiming to use generative AI to “replace as many artists as possible.”
The YouTube caption dataset contained a total of 48,000 channels and 173,536 videos. Some of the videos contained conspiracy theories and parodies, which may affect the integrity of the final model.
This isn't the first time YouTube has been at the center of an AI training data controversy: OpenAI CTO Mira Murati has been unable to confirm or deny whether YouTube was used to train the company's advanced (but as yet unreleased) AI video model, Sora.
Nebula CEO Dave Wiskas told Wired that using their data without their consent is “theft” and “disrespectful,” given that studios are already using generative AI to “replace as many artists as possible.”
In a statement to Ars Technica, Anthropik said that Pile is only a small subset of YouTube's subtitles and that YouTube's terms only cover direct use of the platform, which is different from use of the Pile dataset. “Any concerns about potential violations of YouTube's terms of service should be directed to the Pile authors.”
What will happen to AI?
Google says it has taken steps over the years to prevent abuse, but has not provided further details about what those steps are or whether this violates its terms.
However, Google is not entirely innocent, as it was found to have used Gemini AI to scan user documents stored in Google Drive even without users' permission.
Creators are outraged by the discovery, but issues of origin and copyright of the data used to train the models remain up for debate, and creators' only recourse would be if Google determines they violate YouTube's terms.
This potential misuse of data may be rolled up into the broader issue of whether training data qualifies as fair use or requires a specific license, a matter that I suspect may take years to be finalized.
