Why AI companies are offering free subscriptions in India | Technology News

AI News


When I opened the Airtel app on my smartphone to top up my mobile plan over the weekend, I was surprised to see that the plan included a Perplexity Pro AI subscription worth Rs 17,000. I double-checked to make sure I didn't subscribe to Perplexity AI by mistake.

Before I panicked that I might be subscribing to another service, I noticed this: Perplexity AI Subscription It's free on my plan and the offer is valid until July 2026. But Perplexity is not the only company in India offering free AI subscription plans to millions of people. Another telco, Reliance Jio, has an active partnership with Google and is offering 18 months of free Google AI Pro subscription to some of its users.

OpenAI, on the other hand, offers its annual “low cost.” ChatGPT Go subscriptionworth Rs 399 per month, free for 1 year in the country.

It's not that AI companies are comfortable offering free services to users who buy their products.

For example, Apple is known for offering three months of free access to Apple TV along with new products such as the iPhone and iPad. My new Pixel 10 Pro Fold and 10 Pro XL also come with a free Google AI Pro plan that includes premium AI features and 2TB of storage for 1 year. In both cases, you'll be paying a premium for the hardware product, but these technology companies may offer access to new services to increase their popularity and lock you into their ecosystem.

However, Google, OpenAI, and Perplexity are offering long-term free access to their premium AI services to millions of users in India. One might wonder why these tech companies are so generous when artificial intelligence is an expensive playground and the cost of infrastructure investment and training of AI models runs into billions of dollars.

Well, these tech companies may be showing a calculated motive rather than a naive move to provide premium AI tools for free to millions of Indians considering the long-term rewards are high.

Story continues below this ad

Home to millions of smartphone and internet users

India is the world's largest consumer of mobile data per user and its internet users are expected to exceed 900 million, creating huge market potential. This boom is largely due to the availability of low-cost internet, widespread adoption of smartphones in rural areas, digitally savvy young people aged 18 to 35, and growth in digital infrastructure and services such as digital payments. This is why many companies are pouring money into India's burgeoning internet ecosystem, and for many investors this award is considered too big to ignore.

The role of telcos like Airtel and Jio is equally important in setting the stage for India's AI boom. Both Airtel and Jio have large user bases that play a decisive role in ensuring effective execution from deployment to demand transformation, ultimately driving greater user adoption of any new service.

….But where does the AI ​​training data come from?

No wonder companies like Google and OpenAI see an opportunity to train cutting-edge advanced algorithms and models in India. The problem is that all AI models from GPT to Gemini require vast amounts of human-labeled data, and a country like India could be ideally suited to be the backbone of training data. But the big question is, where will the AI ​​training data come from?

To build large-scale generative AI models, tech companies often rely on the public internet, but there's no single place to download the entire web. Instead, technology companies select training sets using automated tools that catalog and extract data from the internet. After all, high-quality data, primarily collected from the web, is critical to the performance of AI models.

Story continues below this ad

These tools include web “crawlers” called “spiders.” It is an automated program that systematically browses the World Wide Web and indexes its pages. For example, Google's parent company Alphabet has already built a web crawler to power its search engine and can use its own tools to collect data and train AI models. But other companies rely on resources like Common Crawl, the primary source of training data that powers OpenAI's GPT. GPT memorizes large amounts of text to answer user questions when given a prompt.

In addition to publicly available data, AI companies also use their own data to train their models. For example, OpenAI fine-tune models based on user interactions with chatbots. Meta AI is partially trained on public posts from Facebook and Instagram. Amazon also says it uses some voice data from customers' Alexa conversations to train LLM. However, for the most part, AI companies have been secretive about the datasets used for training.

Now comes the difficult part.

Lack of transparency around training data is the biggest red flag that AI companies are under scrutiny for. The New York Times recently filed a lawsuit against Perplexity, alleging that Perplexity illegally copied and distributed copyrighted content.

The lawsuit, filed last week in the Southern District of New York, accuses Perplexity of illegally scraping Times articles, videos, podcasts and other content to create answers to users' questions. Another publication, the Chicago Tribune, filed a similar copyright lawsuit against Perplexity. The Tribune also alleges that Perplexity scraped and distributed its content without permission.

Story continues below this ad

Perplexity, founded by Indian-born Aravind Srinivas, is the subject of multiple lawsuits. Earlier this year, leading digital infrastructure company Cloudflare accused Perplexity of concealing its web crawling activities and scraping websites without its permission. Mr Perplexity denied the charges.

In October, social media company Reddit also sued Perplexity in New York federal court, alleging that it and three other companies illegally scraped data to train its AI-based search engine.

News sites and numerous publications have accused AI companies of using copyrighted content without permission to build and operate their AI systems. In 2023, the New York Times blocked OpenAI's web crawler, GPTBot, from using its content to train AI models. AI companies quickly realized they needed to enter into contracts with publications because the data could not be used properly without permission.

This led OpenAI to begin signing agreements with major international media companies to use their copyrighted content as training data. Axel Springer, France's Le Monde and Spain's Prisa Media signed deals with the ChatGPT maker to provide materials for training AI models, followed by the Financial Times, which will allow ChatGPT users to receive summaries, quotes and links to FT articles.

Story continues below this ad

Reuters and the Associated Press subsequently signed deals with OpenAI, as did Hearst, the Guardian, Condé Nast, Vox, Time, and The Atlantic. Microsoft has signed a deal with USA Today. Meanwhile, Perplexity now has access to articles from AdWeek, Fortune, Stern, The Independent, and Los Angeles Times. Axios, a leading technology publication, also signed a licensing agreement with OpenAI.

However, publishers are fine with search engines like Google using web crawlers to access their websites. In return, search companies receive direct traffic to their content.

Still, conflicts between content creators, publications, musicians, artists, and AI companies continue, with stakeholders going to court to block what they see as infringements on creative rights by AI companies. For example, Disney and Universal recently sued artificial intelligence company Midjourney over an image generator. Two Hollywood studios claim this is a “bottomless pit of plagiarism.”

They claim that Midjourney's tools create “countless” copies of characters such as Darth Vader from Star Wars, Elsa from Frozen, and the Minions from Despicable Me. At the end of the day, transparency regarding data sources should be a priority. Even if an artist or musician signs a deal with an AI company, there's always the question of how well the AI ​​will try to recreate their style with just a few keystrokes. There is no concrete answer yet.

Story continues below this ad

Can Indians protect their data from AI?

Giving away AI subscriptions worth thousands of rupees for free is not a new strategy. Google and others have shown that this strategy has worked in the past and could work again if access to AI services is provided for free. In fact, companies like Google have gained access to large numbers of customers by offering their services for free. For example, consider the company's Google search engine. It is essentially free, but it displays ads on results pages and collects user data, which is where most of its revenue comes from.

However, there is always a catch: making online services free comes at a cost, and ultimately we, the consumers, pay the price. A startup like Perplexity needs just one thing. It's your attention. The goal is to build a sizable user base, and if successful, it can secure funding to grow even further.

Perplexity's valuation increased to $20 billion in just three years. The same goes for OpenAI, which has amassed 800 million weekly ChatGPT users and pushed its valuation to $500 billion. That's how capitalism works.

All major AI companies are eyeing India, and for good reason. India not only has a large consumer base but is also a hub for the outsourcing IT industry. Meanwhile, global AI companies are moving into highly diverse consumer markets where users speak multiple languages ​​and each region has its own culture and dialect, especially in rural towns.

Story continues below this ad

At the same time, we are also gaining access to a large user pool from startups and small businesses. For large technology companies, the more users of their AI services, including students, corporate professionals, and warehouse workers who manage their systems, the better suited they are for training their AI models. And there is no better market than India at the moment.

As the AI ​​ecosystem develops, there is also the potential for a market for cloud farming in small centers across India, which can be used to build datasets to train AI and moderate content.

One fundamental question that cannot be ignored is whether sensitive personal data can be kept away from AI training. Currently, there are no laws in India that specifically regulate artificial intelligence. The Digital Personal Data Protection Act 2023 (DPDP) provides broad protection for personal data but has not yet been enacted. Additionally, the law does not address the liability of AI systems or algorithms.

In states such as California, digital privacy laws give consumers the right to request that companies delete their personal data. In the European Union, artificial intelligence laws impose restrictions on “high-risk systems” used in areas such as education, healthcare, law enforcement and elections. Completely ban the use of some AIs.

Story continues below this ad

However, there is currently no clear way to make AI “forget” previously learned data. To completely remove copyrighted or sensitive information, models must be retrained from scratch, which can cost tens of millions of dollars.





Source link