After animating memes in recent days, the AI is now turning its attention to silent videos — specifically, adding audio to AI-generated clips.
Google's research division, DeepMind, has built a powerful new AI model that can add audio to videos that don't have sound, as well as dub sound effects and music.
What's most impressive about this new research is its ability to accurately track footage: One clip shows a close-up of a guitar being played, and the SFX music matches up remarkably well with what's actually being played.
In a way, this is the flip side of ElevenLabs' experiment last month with generating music based on visual prompts, and it has great potential for restoring older media that's lost its audio element — and, if it goes further, Charlie Chaplin might just get a new voice.
While the Google DeepMind model isn't available yet, you can try out a similar tool from ElevenLabs right now: If you want to create a video and give it a try, check out our list of the 5 best AI video generators.
Google's new audio generation is off to a good start
In the X post thread, Google's DeepMind account starts off with footage of a character walking through an eerily lit tunnel.
Light choral music can be heard over dramatic percussion sounds, as well as footsteps as characters move through the scene.
The second is an audio generated by the prompt “Wolves howling at the moon,” which works well with the animation and even features a chorus of howls in the distance.
We're sharing our progress on Video to Audio (V2A) generation. 🎥Add sound to silent clips that match the acoustics of the scene, or add sound to the on-screen action. Here are 4 examples. Turn on the sound. 🧵🔊 https://t.co/VHpJ2cBr24 pic.twitter.com/S5m159Ye62June 17, 2024
The harmonica example sounds a bit “uncanny valley” in the way the pitch changes, but the backing underneath is solid, while the jellyfish example sounds just like a jellyfish – though notable is the inclusion of additional prompts, such as “marine life” and “ocean.”
But the video, which begins with the prompt “Drummer on stage at a concert, surrounded by flashing lights and a cheering crowd,” is a bit off. For one, the beat doesn't quite line up with the rhythm as the video begins, and the sticks seem to be focused on the snare and possibly the floor tom. Also, the audio sounds a bit convoluted, with other drums involved.
Still, it's an impressive start for a project that will likely grow over time.
Limitations of the DeepMind model
Like many of Google's projects, this isn't released yet and is only a research preview, and Google says there are limitations and safety issues that need to be addressed first.
For example, “Because the quality of the audio output depends on the quality of the video input, artifacts or distortions in a video that are outside the model's training distribution can result in noticeable degradation of audio quality.”
We are also working on lip syncing videos with audio, which we are currently attempting, but it's not always accurate and creates an uncanny valley effect.
ElevenLabs is also working on a similar project.
We're excited to introduce our text-to-sound effects API. To showcase it, we've built the first video-to-sound effects app, which is freely available online and fully open source. pic.twitter.com/8aalo8GCSoJune 17, 2024
Not to be outdone, ElevenLabs this week announced a new Text to Sound Effects API that lets you generate audio effects based on content you upload.
Unlike Google's V2A model, ElevenLabs' API is already accessible, and it's worked surprisingly well in experiments.
In the example above, the bottle-breaking video has a few options available, and the meme of DiCaprio laughing has additional audio from others in the room.
To demonstrate what the API can do, the company has “bootstrapped” a simple app that lets you upload video and add sound. It's free to use, open source, and you can try it out right now.
ElevenLabs told Tom's Guide that its real aim is to enable other companies and developers to use the API to build things like generative video integrations, for example.
