Skyra enables AI video detection using grounded inference and new 4K ViF-CoT dataset

AI Video & Visuals


As artificial intelligence-generated videos become more prevalent, identifying authentic content from synthetic media becomes an increasingly challenging task, requiring robust detection methods. Tsinghua University's Yifei Li, Wenzhao Zheng, and Yanran Zhang, along with their colleagues, are addressing this important need with Skyra, a new system that identifies obvious visual inconsistencies, or artifacts, in AI-generated videos. Unlike existing approaches that focus solely on classifying videos as real or fake, Skyra actively reasons about these grounded artifacts, providing both accurate detection and easy-to-understand explanations for its decisions. The team developed a large, meticulously annotated dataset, ViF-CoT-4K, and a new training strategy to enhance Skyra's ability to recognize these subtle temporal discrepancies, ultimately achieving superior performance against existing methods and providing valuable insights into the development of explainable AI for media authentication.

Video reliability analysis with visual discrepancies

Scientists have developed Skyra, a system that can analyze videos and determine their authenticity by identifying certain visual inconsistencies and artifacts. Skyra acts as an analysis tool that classifies these inconsistencies, such as shape distortions and camera movement inconsistencies, classifies videos as fake or real, and determines authenticity based on observed characteristics. Researchers are now ready to receive and analyze video content descriptions, frame by frame or as summaries, and provide authenticity judgments along with the types of artifacts detected (if any). The analysis follows a clear format. This means describing the observed characteristics of the video, identifying the types of artifacts detected, and ultimately determining whether the video is fake or real.

Skyra model and ViF-CoT-4K dataset development

The research team developed Skyra, a specialized multimodal large-scale language model, to reliably detect and, importantly, explain AI-generated videos. To make this possible, scientists built ViF-CoT-4K. ViF-CoT-4K is a new large-scale dataset specifically designed for supervised fine-tuning and is the first resource of its kind to provide detailed human annotation of AI-generated video artifacts. This dataset consists of high-quality samples and forms the basis for training a model that can identify inconsistencies commonly present in synthetic videos. The team then implemented a two-step training strategy to systematically enhance the model's ability to recognize spatiotemporal artifacts, clarify its inferences, and ultimately improve detection accuracy. To rigorously evaluate Skyra's performance, researchers introduced ViF-Bench, a benchmark consisting of 3,000 high-quality video samples produced by over 10 state-of-the-art video generators to ensure comprehensive evaluation across a wide variety of synthetic content. Experiments demonstrate that Skyra outperforms existing methods across multiple benchmarks, achieving superior performance in both detection and explainability, and providing valuable insights to advance the field of explainable AI-generated video detection.

Skyra detects and explains AI-generated videos

Scientists have developed Skyra, a new multimodal large-scale language model (MLLM) designed to identify and reason about artificially generated videos, addressing a critical need in the face of increasingly realistic synthetic media. To facilitate the training and evaluation of this system, the research team built ViF-CoT-4K, a large dataset of AI-generated videos with detailed human annotations. This dataset represents a significant advance and provides the first resource of its kind for supervised fine-tuning of AI-generated video artifact detection. The team employed a two-step training strategy to enhance Skyra's ability to recognize subtle spatiotemporal artifacts, provide clear explanations, and improve detection accuracy.

Initial supervised fine-tuning of ViF-CoT-4K gave the model significant detection and explanation capabilities and prepared it for subsequent reinforcement learning. This initial stage involved full parameter fine-tuning of Qwen2.5-VL-7B using a learning rate of 1e-5 in 5 epochs. Further improvements were achieved through reinforcement learning using group-relative policy optimization algorithms. The reward function encourages the model to actively explore potential counterfeit clues while strictly adhering to the prescribed output format.

Reflecting the inherent difficulty in comprehensively verifying the authenticity of real-world videos, accuracy rewards were designed asymmetrically, with harsher penalties for false positives. This asymmetric design prevents overfitting and allows the model to focus on identifying even a single artifact as evidence of manipulation. Extensive testing on ViF-Bench demonstrates that Skyra outperforms existing methods in both detection accuracy and explainability, providing valuable insights for advancing explainable AI-generated video detection. The researchers observed that by encouraging active cue exploration and strict adherence to the prescribed output format, the model's performance improved significantly, demonstrating the importance of both accuracy and explainability in this important area.

Accurately identify AI video artifacts with Skyra

Skyra represents a significant advance in the detection of AI-generated videos, moving beyond simple identification to identifying and explaining visual discrepancies that reveal their artificial origins. Researchers have developed a new multimodal large-scale language model, Skyra, specifically designed to identify human-perceptible artifacts in videos created by artificial intelligence. This approach not only enables accurate detection of fake videos, but also provides grounded evidence in the form of local visual anomalies, providing a level of transparency not present with existing binary classification methods. To facilitate this achievement, the team built ViF-CoT-4K, a large dataset of AI-generated videos annotated with detailed information about the artifacts present, enhancing Skyra's ability to recognize subtle production flaws.

Further improvements through reinforcement learning increased the model's ability to identify artifacts, resulting in significant increases in both detection accuracy and explainability. Extensive evaluation using the ViF-Bench benchmark demonstrates that Skyra consistently outperforms existing methods. The authors acknowledge that current methods still face challenges when it comes to very high-quality AI-generated videos with minimal or no perceptible artifacts. Future research directions include exploring ways to improve robustness to increasingly sophisticated generative models and extending the dataset to cover a wider range of video types and artifact characteristics. Despite these limitations, Skyra establishes a promising new direction for explainable AI-generated video detection, providing an important tool for combating the spread of misinformation and maintaining trust in visual media.



Source link