Microsoft this week demonstrated VASA–1, a framework that creates videos of people talking from still images, audio samples, and text scripts, and rightly claimed it was too dangerous to release to the public.
These AI-generated videos are convincingly animated with people speaking scripted words in replicated voices, and the U.S. Federal Trade Commission last month passed rules that prevent the use of AI technology. This is exactly the same thing you warned about after suggesting . In case of identity fraud.
The Microsoft team acknowledged as much in their announcement, explaining that the technology will not be released due to ethical considerations. They claim that they are presenting research to generate virtual interactive characters, not to impersonate anyone. Therefore, there are no plans for a product or API.
“Our research focuses on generating visual-emotional skills for virtual AI avatars and aims for positive applications,” Redmond officials said. “It is not our intention to create content that will be used to mislead or deceive.
“However, as with any related content generation technology, there is still the potential for it to be misused to impersonate humans. We disagree and are interested in applying our technology to advance counterfeit detection.”
Kevin Surace, chairman of biometrics industry Token and frequent speaker on generative AI, said: register In an email, Microsoft said that while there have been previous demonstrations of technology that animates faces from still frames and cloned audio files, Microsoft's demonstration reflects cutting-edge technology.
“The implications for personalizing email and other business mass communications are impressive,” he said. “You can also animate old photos. In one sense, this is just fun, but in another sense, it's a solid business application that we'll all be using in the coming months and years. ”
When evaluating the “fun'' of deepfakes, 96% said they were non-consensual pornography. [PDF] It was announced by cybersecurity company Deeptrace in 2019.
Nevertheless, Microsoft researchers suggest that there are positive uses for being able to create realistic-looking people and type words into their mouths.
“Such technologies will enrich digital communication, increase accessibility for people with communication disorders, transform education, provide ways to utilize interactive AI tutoring, and improve therapeutic support and social interaction in healthcare. ”, they propose in a research paper. The word “pornography” or “misinformation.”
While it's arguable that AI-generated videos are not exactly the same as deepfakes, which are defined by digital manipulation rather than production techniques, they can be convincingly faked without cut-and-paste grafting. If we can produce , the distinction becomes unimportant.
When asked what he thought about the fact that Microsoft is not releasing this technology to the public for fear of misuse, Surace expressed doubts about the feasibility of the restrictions.
“Microsoft and others are holding off for now until privacy and usage issues are resolved,” he said. “How do we regulate people who use this for legitimate reasons?”
Surace added that similarly sophisticated open source models already exist, pointing to EMO. “You can take the source code from GitHub and build a service around it that probably rivals Microsoft's output,” he said. “Since this field is open source, it is impossible to regulate it in any way.”
However, countries around the world are trying to regulate people created by AI. Canada, China, and the United Kingdom all have regulations applicable to deepfakes, some of which serve broader political goals. The UK this week made it illegal to create sexually explicit deepfake images without consent. Sharing such images is already prohibited under the UK's Online Safety Act 2023.
In January, a bipartisan group of U.S. lawmakers introduced the Defending Explicit False Images and Nonconsensual Editing Act of 2024 (DEFIANCE Act). This bill would create a way for victims of non-consensual deepfake images to file civil lawsuits in court.
And on Tuesday, April 16, the U.S. Senate Judiciary Committee's Privacy, Technology, and Law Subcommittee held a hearing entitled “AI Surveillance: Election Deepfakes.”
Rijul Gupta, CEO of Deepmedia, a deepfake detection business, said in prepared remarks:
But let's think about marketing applications. ®
