video Mainstream adoption of generative AI technologies has focused primarily on the creation of text and images, but ultimately the statistical probabilities on which these models are based are just as valid for generating any other kind of media.
The latest example of this came on Monday, when Google's AI lab DeepMind published details of its work on a video-audio model that can generate sound to match video samples.
The model works by taking a video stream and encoding it into a compressed representation, which, along with natural language prompts, acts as a guide for the diffusion model, which goes through multiple steps to refine the random noise into something that resembles the audio associated with the input footage. This audio is then converted into a waveform and combined with the original video source.
From our understanding, this approach is not very different from how image generation models work, but instead of outputting photos or illustrations, they are trained to reproduce audio patterns from video or text inputs.
Below is one of several examples DeepMind released this week that shows the model in action.
DeepMind says they used a variety of datasets, including video and audio, naturally, but also AI-generated annotations and transcripts to help teach the model to associate different visual events with different sounds. As the researchers explain, this means the model can generate audio with or without text prompts, without the need to manually tune the tracks. However, there are still some hurdles to overcome.
Firstly, due to the way the audio is produced, the actual quality of the soundtrack depends on the source material – if the video is poor quality, the audio is likely to be poor as well. Lip syncing also proves to be rather difficult to say the least.
DeepMind hopes that this new model will pair well with models designed for video generation, including its home-grown Veo model.
According to the DeepMind team, one problem with current text-to-video conversion models is that they are typically limited to generating silent films. By combining them with their video-to-audio conversion model, the team claims that fully AI-generated videos, complete with soundtracks and even dialogue, are possible.
Speaking of video generation models, the sector has grown significantly over the past year, with more players entering the space.
ML giant OpenAI unveiled its own video generation model, called Sora, in February. But Sora is just one of several models that push the boundaries of what's possible.
Cling AI
Among these models is that of Kling AI. Developed by Chinese tech company (partly state-owned) Kuaishou, Kling uses a combination of diffusion transformers to generate frames and a “3D spatiotemporal attention system” to model movement and physical interactions in a scene. Here's what the system looks like in action:
The resulting videos could easily be mistaken for human-created video footage if you didn't look closely. But look closer and you'll quickly start to notice visual artifacts and inconsistencies. That said, this seems to be a common theme among many of the AI video generators on the market today.
Few details are known about Kling, but its developers claim that it outperforms OpenAI's Sora. The model can produce up to two minutes of video at 1080p resolution and 30 frames per second. Unfortunately, access to the model is currently limited to China.
Runway
Another model builder working on video generation is Runway, which rolled out its Gen-3 Alpha model on Monday. Runway has been working on a number of image and video generation models since early 2023.
Runway says Gen-3 Alpha is one of several models currently in development that was trained on a combination of video and imagery with highly detailed captions. The startup says this allows it to deliver more immersive transitions and camera movements than previous models. Here's the model in action:
Embedded MP4 Video
The model will also reportedly introduce safeguards to prevent users from generating inappropriate images and videos. Runway plans to incorporate the latest models into its existing library of text-to-video, image-to-video and text-to-image services, as well as work with industry partners to create custom versions.
Pika
These AI video startups are clearly catching the attention of investors: Earlier this month, Pika raised $80 million in Series B funding from Spark Capital and others to accelerate the development of its AI video generation and editing platform.
The team's 1.0 model, released in beta last July, supports a variety of common scenarios, such as generating videos based on text, image, or video prompts. Over the past few months, Pika has added support for fine-grained editing, inpainting style adjustments, sound effects, and lip sync. Here's what it currently has:
Similar to the popular generative image model Midjourney, users can interact with Pika's AI video service through Discord or the startup's web app.
Of course, this is not a complete list. As AI developer ambitions move beyond text and image models, we expect to see many more models and video generation services emerge in the coming months. If you've seen anything related to this, please let us know below.®
