As deepfakes continue to evolve, the boundary between real and fake videos becomes difficult to detect. Today's synthetic content goes far beyond face swaps and fake lip syncing. Thanks to powerful generation tools, the entire scene, such as background, lighting, and movement, can now be manufactured from scratch. And as these tools become more accessible, the more likely it is to be harmed.
In response to this challenge, researchers at the University of California, Riverside collaborated with Google scientists to develop a groundbreaking system. Model known as Unite – Abbreviation Universal Network for Identifying Tampered Composite Videos – It is built to catch deep fakes in all forms, not just those that include faces.
Why fake videos are becoming harder to find
Many deepfake detectors work by looking closely at your face. They look for strange flashing patterns, inconsistencies in the lighting, or unnatural movements around the mouth and eyes. But that's not enough.
“Deepfake has evolved,” said Rohit Kundu, a doctoral candidate in computer engineering at UC Riverside. “They're not just face swaps anymore. People are now using powerful generative models to create completely fake videos from face to background. Our system is built to catch all of that.”
Kundu has worked with advisors, Professor Amit Roy-Chowdhury, and researchers at Google to develop the solution. Their models focus not only on the face but on the entire video frame. This includes movement, textures and even background, making it one of the first tools to detect tampers at a wider visual level.
“It's scary how accessible these tools have become,” Kundu added. “Anyone with moderate skills can bypass the safety filter and generate realistic videos of public figures saying things they never said.”
Tools like text-to-video (T2V) and image-to-image (I2V) generation make this even easier. These technologies use artificial intelligence to convert written text or still photos into realistic video clips. And with these clips present it makes it difficult to know what is being staged and what is real.
Related Stories
How Unite Model Works
Unite systems take a different approach than older detectors. It uses a deep learning method called trance. It processes data sequences, such as video frames, while tracking both spatial and temporal patterns.
Instead of focusing on human faces, we look at domain-independent features. This means that the system does not depend on who or what is displayed in the video. Instead, we look at more general characteristics, subtle details of movement, color shifts, and object placement. Unite builds its backbone on a basic model called the Siglip-SO400M. This powerful AI model processes large amounts of visual and linguistic data. Even clips without human subjects can help analyse a wide range of content.
Another innovation lies in Unite's training strategy. Many AI models struggle because they concentrate so heavily on the most obvious cues, such as the face of a person. However, Unite uses a new loss function called loss of attention. This feature teaches a system to examine different parts of each video frame, spreading attention throughout the scene. “It's one model that handles all these scenarios,” says Kundu. “That's the Universal thing.” This combination of features allows Unite to detect video tampering even when people are not visible, such as empty rooms, modified environments, and Ai-generated clips in animation landscapes.
Tools for ages of disinformation
The team shared the success of the project at a 2025 conference on Computer Vision and Pattern Recognition (CVPR) in Nashville. Their paper explains how unity works and what stands out from other systems.
This study includes contributions from Kundu, Roy-Chowdhury and three Google scientists, Hao Xiong, Vishal Mohanty and Athula Balachandra. Kundu's internship at Google provided access to vast training data and computing resources to support research.
Instead of training only on standard deep-fark datasets, the team united into a variety of content types. This included data relating to tasks such as unrelated video footage to reduce the risk of overfitting. As a result, Unite performs better in real life situations, not just in laboratory testing.
To test accuracy, researchers evaluated Unite using several benchmark datasets. These featured facial manipulation, background changes, and a complete composite video. In almost every case, Unite was better than the best existing detectors. “People deserve to know if what they are seeing is real or not,” Kundu said. “And as AI gets better at fake reality, we have to make it better for revealing the truth.”
What's next for synthetic video detection?
As fake videos become more persuasive, they can cause serious harm. Bad actors use them to spread false political messages, harass individuals, and promote dangerous hoaxes. Viral clips often spread faster than a team can fact check. So tools like Unite can quickly make a huge difference. Although still in the research phase, this model can serve platforms and professionals working in the real world.
Social media companies can use it to scan uploads for signs of tampering. Fact checkers and journalists may rely on unity to confirm the reliability of viral videos. Even law enforcement and government agencies can benefit from more reliable methods of detecting synthetic content.
Roy-Chowdhury, who co-oversees the UC Riverside Artificial Intelligence Research and Education (Raise) Institute, highlights the need for more powerful tools in today's digital world. “This work brings us closer to a built system that can protect harmful video-based misinformation,” he said.
By examining the big picture, not just the face, Unite offers a more flexible and advanced approach. As Deepfake technology continues to grow in its power and reach, these tools may become essential.
Findings are available online via ARXIV.
