Companies like Reddit, Slack, Google, Facebook, and Instagram are using our data directly or indirectly to train the next generation of AI language models. But no one seems to have asked our permission. In doing so, these companies have proven the adage that their customers' data is their primary product.
For much of the internet generation, companies have offered products for free or at little cost to lure customers into their ecosystems. Products like Gmail, YouTube, Facebook, and Reddit appear to be free, but they collect user data that can be used to serve ads or sell it in aggregated bundles.
While these business models were once acceptable, rapid advances in AI have raised much larger and more pressing questions that will have significant implications for the future of privacy.
Understanding AI and LLM
The current generation of AI is based on LLMs (Large Language Models) that recognize, understand and generate human language. Built using machine learning and trained on huge datasets, they can generate human-like text, recognize images, answer questions and process voice and video in real time.
An LLM consists of three main parts: parameters, weights, and tokens. Parameters form the variables that the model learns during the training process. Weights determine the strength of the connections between variables. Tokens form the basic input and output, i.e. the natural language text, audio, or video that you feed into an LLM and receive as a response.
Consider a chef as an example: a customer orders a particular dish (the input token), and the chef puts a set of ingredients into a frying pan to make the dish. The final dish is the output token, while the specific combination of ingredients used to make it are the parameters and the specific recipe represents the weight. All chefs can make that dish (let's assume it's a very basic dish), but the degree to which they can do so varies depending on their knowledge, training and experience.
What is generative AI?
Agents of human will, amplifiers of human cognition. Discover the power of generative AI.
Let's say someone asks Gemini or ChatGPT-4o for a recipe. An LLM can only learn this based on the dataset. The more recipes it ingests (corresponding to the number of times a chef has made that dish), the better it will be at predicting how to make a tasty dish. As a result, the best LLMs will give the best recommendations, especially when asked for a recipe given multiple ingredients.
AI issues loom
The biggest problem with the above is the sheer amount of data required to train an LLM. Here are some examples: OpenAI trained GPT-4 (not the latest model, the latest is GPT-4o) using 1 million hours of YouTube video data. Google DeepMind trained their Gemini model using around 10 trillion words collected from the web. Meta trained their generative AI model using images, videos and text uploaded by users to their platform.
But it doesn't end there. Google paid Reddit $60 million to scrape all of Reddit for its AI. This quickly led to Reddit becoming one of the main sources for its AI Overview feature. But to Google's detriment, the AI lost the battle of AI vs. human internet users by a large margin. Just ask anyone who searches “glue pizza” or “how to eat a stone” on Google.
That money went to Reddit, presumably to be spent since the word Reddit is often followed by many of the most popular search terms when users are looking for human answers, but Reddit's millions of users never see any of it, which is especially odd considering it's those same users who have worked unpaid to build a platform that Reddit can monetize and leverage.
Five big announcements from Google I/O: Circle to Search, search changes, and more AI
This Isn't Your Dad's Google
Reddit is just one example of companies misusing users' data. Meta owns the world's largest platforms, including Facebook, Instagram, and WhatsApp. Elon Musk trains X AI's GrokAI on Twitter, one of the largest sources of real-time information. None of these companies are paying users for this. And many of them are encouraging users to sign up for subscriptions, which means users are paying to provide these companies with their data. But none of these subscriptions allow you to opt out of using your data.
You could argue that all these platforms are free and your data is freely available, and while I somewhat agree if you are not paying for the platform, what if you are and you are still a product?
This is where the line needs to be drawn. What's the inspiration behind this article? Slack is a business-facing service that requires a paid subscription for many of its core features, and it trains its AI using company data, much of which is likely quite sensitive.
When is enough enough?
This raises further questions: when do we say “enough is enough”? We have already seen Google Gemini create AI teammates, created under the guise of reducing friction and communication between different teams, but it is easy to imagine them evolving into replacing full-time jobs. Google's AI Overviews also disrupts the role of journalists and fact-checkers, but as the lawsuits by a number of publishers suggest, this started in other business practices at Google long ago.
Companies using our data for their own gain without compensating users is nothing new. Lou Monturi created the digital cookie in 1994, and within a year, targeted advertising to specific consumer demographics became the norm. For over 20 years, digital customer privacy has not been a priority, and without GDPR (a 2018 EU ruling), we probably wouldn't have a concept of privacy yet. Instead, we now have companies monetizing user data by taking everything you post on the web and training AI on it.
AI will inevitably change our digital lives, but not necessarily for the better. Companies like OpenAI have deals with big publishers (with big budgets) like Vox Media, but most people will not benefit from it. Rather, everyday users will still be the commodity. The solution seems simple: find a way to compensate users. Given that Google, Meta, and others have threatened to stop distributing content in certain states and countries to avoid paying publishers, it is highly unlikely that companies will pay for users' data. So, if we are not going to be compensated for the knowledge these multinational companies are using to make huge profits, as the headline of this article says, companies need to stop using our personal data to train AI. Because if we continue on our current path, the only people who can create the free content/data we consume will be the companies that stole ours.
