Image credit: Bryce Durbin/TechCrunch
Services like Midjourney and ChatGPT are pushing the boundaries of how AI can create images and text from basic text prompts. Audio now seems to be the inevitable next frontier. Music generation based on word prompts, his AI tutor for language learning, and voice simulators have all made progress in recent months. Voice.ai wants to engage in that conversation (huh) with technology that allows users to change (and impersonate) their voice in real-time, and is now raising its first external funding following its initial growth. procured.
With over 480,000 users and a library of over 50,000 voice filters, Voice.ai has raised $6 million and plans to use the funds to bring voice modification technology into new areas.
Mucker Capital and M13 are leading the round. To date, Voice.ai has grown by word of mouth, backed by his $3 million of self-funding. The startup has a Discord channel with over 120,000 members.
Currently available as apps for Mac, PC, Android, and iOS, the company’s tools are used by TikTok, Zoom, Discord, Minecraft, GTA5, Fortnite, Valorant, League of Legends, Among Us gamers, content creators, Vtubers, and more. I’m here. , Skype, Whatsapp and other platforms. Using the Voice.ai interface, you can create new voices or choose from nearly 50,000 different pre-made voices (created and shared by users like yourself). These voices can be used as is or modified and used live on any supported platform. for recording.
The plan is to use this funding to hire more tech talent and build new SDKs and APIs to work with additional platforms such as Meta, Unreal and Unity. Introduce multilingual support. In addition, new applications such as singing where the voice is the star can be added.
The startup isn’t specifically opting for it, but it will be interesting to see if it also uses some of the funding to increase server capacity.
It is no small burden. As an aside, I’ve heard that GPU pain is one of the biggest factors in how many AI apps can scale right now. (This is one of the reasons why big deals are being made involving strategies that provide processing power and server capacity.)
For Voice.ai in particular, your voice is processed locally and channeled to where it’s used through what founder and CEO Heath Ahrens described to me as a “virtual audio cable.” But looking at reviews for that app, a common lament is that “overwhelming demand is pushing our servers to max capacity”, so you’ll be put on a waiting list when you sign up, and when service ramps up. The promise is that you will be notified. capacity.
There are currently dozens of text-to-speech and text-to-speech services on the market, and there is already a lot of activity going on between them. Last year, Spotify acquired his Sonantic, and Snap acquired his AI voice assistant earlier. Another startup, Sanas, is working on accent modification, including speech simulators His Murf and His Acapela. Voice.ai, which falls into the same general category as text-to-speech AI startups Respeecher and Celebrities, allows users to apply masks to fine-tune their voice or completely transform it. to In some cases, a fully synthesized voice is created instead of the real voice.
Founded and based in Ukraine, Respeach will build a new Darth Vader voice in the new Star Wars movie, based on what it sounded like 45 years ago when James Earl Jones conceived the role. He made a name for himself by helping the (According to a character bent on destroying the world, Darth’s voice was delivered from his Ukrainian office to a client in Hollywood as Russia marched into the country.)
Eleven Labs, famously (and in some cases notoriously) building a platform that is terrifyingly good at cloning voices, received $1900 in its latest funding round from a group of high-profile investors earlier this month. raised millions of dollars.
Voice.ai is trying to establish itself as an AI voice modification app for Everyman in that situation.
“There are a lot of companies trying to offer different voice technologies to companies,” Ahrens told TechCrunch via email (ironically, I was unable to arrange a live interview with him. ). Ahrens has some experience building B2B AI technology. His two previous companies, his iSpeech for text-to-speech and his Haystack for facial recognition, are built around API services.
“Voice.ai is unique in that we are focused on bringing technology that has traditionally been available to businesses, directly into the hands of consumers at an affordable price.” said, “It comes to us from the classic DSP voice changers and voice modulators that we have used in the past and are still popular among many gamers and streamers.”
There are two tiers of “affordability”, and most users currently have a free service that requires them to opt-in to provide computational power to train Voice.ai’s models. Its service is “built on a unique private data set consisting of millions of unique data sets for ‘users'”. I am asking for details as the site does not list any prices.
“We believe in making technology accessible, and we plan to work with the open source community to democratize voice AI technology,” Ahrens added.
Voice.ai also leverages some of the ethos that has been built around Vtubers, gamers, and other online avatar uses, taking a radically different approach to the challenge of changing voices. claim to be.
“Most of the voice AI companies trying to enter the space are trying to build scalable enterprise text-to-speech solutions or expensive text-to-speech services for production studios,” said Ahrens. “We are starting from the opposite spectrum and looking to deliver value to individuals looking to expand their sound online. It’s not that a person can be perfectly replicated, it’s that the tone of the voice is preserved while retaining the core elements of the user’s speech: emotion, pacing, and emphasis, to create a completely unique new end result in real time. to replace.”
Maybe it’s because of the skewed demographics of interactive platforms like games, but for now, Voice.ai’s audience is 70% male and 30% female, and who are using the technology? It opens up a new category of why you are using the technology, not just if it is.
This includes not only users who use avatars and create voices to match them, or who seek more privacy protection, but also “transgender users who can express themselves with voices that match their identity, This includes users exploring a whole new online world,” he said. A persona for yourself. ”
While there is already a user base leveraging Voice.ai’s direct-to-consumer service, one of the reasons Mucker invests in the startup is the opportunity to build a network of developers who use and integrate with Voice.ai. Because we think there is. that technology.
“Voice.ai aims to revolutionize the AI developer community in the same way that AdMob has impacted the mobile app developer community,” said Omar Hamoui, partner at lead investor Mucker Capital. Told. (Hamoui previously founded mobile advertising startup Admob, which was eventually acquired by Google, so he has direct experience building mobile developer tools.) Developers around the world. ”
Karl Alomar, former COO of Digital Ocean, which led the M13 investment, said investors would play an active role in the next phase of development. “Even at Digital Ocean, we recognized the value of building a community of builders by builders,” he said. “We look forward to what creators and developers can build on his Voice.ai platform.”
