Microsoft shows off generative AI model that creates deepfake videos from still photos

AI Video & Visuals



A team of AI researchers at Microsoft Research Asia has demonstrated a new generative AI application that can generate deepfake videos of people from just static images. The new VASA-1 model accurately depicts an individual speaking or singing, creating animations synchronized to audio tracks with appropriate facial expressions.

VASA-1

VASA-1 is named for its ability to create videos using visual affective skills (VAS). The researchers aimed to animate still images with realistic expressions that synced seamlessly with a provided audio track. VASA-1 was trained on thousands of images depicting different facial expressions. Through extensive experimentation and development, they were able to create animations that were synchronized enough to fool casual viewers. Of course, upon closer inspection, the researchers recognized the existence of flaws indicative of artificial generation.

Still, the effectiveness of VASA-1 has been demonstrated through various video samples shared by the research team, demonstrating its ability to animate subjects as diverse as cartoon characters, personal photos, and even hand-drawn images. I am. In each case, facial expressions dynamically change depending on the words spoken or sung, increasing the sense of realism throughout the animation. Below are some examples of videos expressing different emotions.

“Our premier model, VASA-1, not only produces lip movements that are exquisitely synchronized with the audio, but also captures a wide range of facial nuances and natural head movements that contribute to the perception of authenticity and liveliness. The core innovations include a generative model of global facial dynamics and head movements that works in the facial latent space, and the use of video to create such expressive and disentangled models. “It involves the development of a facial latent space,” the researchers explain in their paper. “It paves the way for real-time engagement with lifelike avatars that emulate human conversational behavior.”

VASA-1 has potential applications in generating lifelike avatars for games and simulations, especially when combined with other deepfake technologies like the VALL-E synthetic voice clone model that Microsoft introduced last year. Masu. The research team warned against using VASA-1 too quickly due to concerns about potential misuse and ethical implications. As a result, the system is currently not publicly available, reflecting the researchers' commitment to responsible AI development.

Microsoft announces 3-second voice cloning tool “VALL-E”

Pindrop launches real-time audio deepfake detection tool Pindrop Pulse

Generated AI video startup D-ID partners with deepfake audio startup Eleven Labs








Source link

Leave a Reply

Your email address will not be published. Required fields are marked *