A new update to Google's video-to-audio tech will allow users to apply AI-generated scores, sound effects, and dialogue to video clips.
According to a blog post from the tech giant, the technology “combines video pixels with natural language text prompts to generate a rich soundscape for the on-screen action.”
Google adds that the update is a “big step towards bringing generated movies to life.”
Read next: New music search engine cosine.club suggests tracks based on similarity
Google's AI research arm DeepMind said the text prompts aren't necessarily necessary because the technology can only “understand raw pixels,” but they can help make the software more accurate.
The technology also comes with increased creative control, with “V2A being able to generate an unlimited number of soundtracks for any video input.”
“Optionally, you can define 'positive prompts' to steer the generated output towards desirable sounds, and 'negative prompts' to steer it away from undesirable sounds,” a Google spokesperson said.
The update has not yet been made public, but the statement added: “However, there are a number of other limitations that we are working to address and further research is ongoing.”
“Because the quality of the audio output depends on the quality of the video input, artifacts or distortions in the video that fall outside the model's training distribution can lead to a noticeable degradation in audio quality.”
Read this next: The Rise of AI Music: A Force for Good or a New Low in Artistic Creativity?
Check out the example clip below to see this technology in action:
Jamal Johnson is a Digital Intern at Mixmag. Instagram
