Microsoft announced its new VASA 1 AI model, a framework designed to generate life-like conversational faces for virtual characters that boast compelling Visual Affective Skills (VAS). The company says VASA-1 can create short, life-like videos with just one still image and one audio clip. This model also offers several options for making changes to your video. Here's everything you need to know.
In a post on its research announcement page, Microsoft revealed a new AI model that can synchronize lip movements with audio and capture a wide range of facial nuances and natural head movements. VASA 1 is claimed to be able to deliver high quality video quality content with realistic face and head dynamics. This model supports online generation of 512 x 512 videos at up to 40fps, ignoring start delays.
It is possible to create a video of up to 1 minute according to the content posted on the official website. As you can see in the embedded video below, the AI model gives users more control, allowing them to adjust various aspects of the video, such as primary gaze direction, head distance, and emotional offset. By controlling the disentangled look, 3D head pose, and facial dynamics, anyone can better modify the output.
(Video credit: Microsoft)
Microsoft's new AI model also shows the ability to process photo and audio inputs that aren't included in the training distribution. For example, it can process artistic photos, song audio, and non-English speech. These types of data were not present in the training set.
However, Microsoft has announced that it will not release VASA-1 to the public, emphasizing its intention to use the technology to create realistic virtual characters, rather than releasing it as a standalone product or API. This decision grew out of his Microsoft commitment to ethical AI practices.
Addressing concerns about the potential for abuse, Microsoft has clarified its position: “Our research focuses on proactive applications that generate visual emotional skills for virtual AI avatars. We oppose any use of this technology to mislead or deceive. Our methods could potentially be exploited for identity theft, and we do not accept any such risks. We are dedicated to advancing counterfeit detection technology to help alleviate this.”
Microsoft does not plan to release online demos, APIs, or additional implementation details related to VASA until Microsoft is confident that this technology will be used responsibly and in accordance with regulations.
