Google DeepMind announces “AI that generates music to perfectly match videos”

AI Video & Visuals



Following AI that generates images and text, video generation AI is also rapidly advancing, but until now, videos generated by AI have only been silent or have human voice added. On June 17, 2024, Google DeepMind announced “Video to Audio (V2A)”, which generates music and audio according to the atmosphere and movement of the video.

Generating Audio for Video – Google DeepMind

https://deepmind.google/discover/blog/generating-audio-for-video/

The V2A system announced by Google DeepMind is a video generation AI called ” Veo .

For example, the following film is accompanied by music and sounds with the prompt: “movie, thriller, horror film, music, tension, atmosphere, footsteps on concrete.”

V2A Horror – YouTube

In the scene where a person is walking back and forth, ominous background music and the sound of crunching footsteps can be heard.

The scene changed and a figure appeared, followed by a heavy “boom” sound.

There are lots of other examples too. The audio prompt for this movie is “Cute baby dinosaur sounds, jungle ambiance, egg cracking sounds.”

V2A Dinosaur – YouTube

“Jellyfish pulsating underwater, marine life, ocean”

V2A Jellyfish – YouTube

“A drummer on a concert stage surrounded by flashing lights and a cheering crowd”

V2A Drum – YouTube

“A car skids, the engine starts, and angelic electronic music plays.”

V2A Cars – YouTube

“A slow, soft harmonica plays as the sun sets over the prairie.”

V2A Cowboy – YouTube

“Wolf Howling at the Moon”

V2A Wolf – YouTube

The V2A system first encodes the input video, then uses a diffusion model to generate repeating sounds from random noise, and once a lifelike voice that matches the video and prompts is generated, it decodes it and synthesizes the audio data with the video.

Because the V2A system can understand video, entering text prompts is optional. For example, the guitar sounds in the video below were synthesized without any prompts.

V2A Guitar – YouTube

It's still prone to being unnatural, but some degree of lip syncing is possible. For example, the lines spoken by the characters in the video below were synthesized from a script: “Transcript: 'This turkey looks amazing. It makes me so hungry.'”

V2A Clay Animation Family – YouTube

In addition to videos generated by Veo, audio can be added to a variety of existing footage, with Google DeepMind saying that “it can also generate audio for a variety of existing footage, such as archival material or silent films, opening up a wide range of creative opportunities.”





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *