AI video generation models have improved dramatically since we released Gen-2, the first publicly available text-to-video conversion model, in early 2023. Two years ago, these models took minutes to produce choppy, pixelated clips that were seconds long. Today’s leading video generation models can reliably produce output that is virtually indistinguishable from real video.
This week, we released image-to-video capabilities for our latest base model, Gen-4.5. Today, we’re announcing new research that evaluates people’s ability to determine whether a 5-second video is real or generated by our model. We’re also launching a new site where anyone can try it out for themselves.
This research study recruited a random sample of 1,043 participants. Each participant watched 20 videos (10 real and 10 generated) in random order and decided whether each was real or generated by an AI. Each video was generated only once, the output was not edited, and the videos were not regenerated to improve the quality or distortion of the results.
result
More than 90% of participants could not reliably distinguish between the Gen-4.5 output and the actual video.
Only 99 of 1,043 participants (9.5%) achieved statistically significant accuracy (accuracy rate 15/20 or higher, p < 0.05, binomial test). The overall detection accuracy was 57.1%, just above chance. The performance is similar on real videos (58.0%) and generated videos (56.1%), indicating the lack of a systematic detection strategy.
Detection accuracy varies by content category. Human-related videos (faces, hands, actions) were more detectable (58-65%), while animals and buildings were less detectable (45-47%). Participants were more likely to mistake the generated video for the real thing than vice versa.
These findings represent a fundamental shift in thinking about video authenticity. For many years, we have been building toward a general world model. Realistic simulation is a prerequisite for solving difficult problems in the physical world. Gen-4.5 is the most powerful simulator we’ve ever built. However, with that ability comes responsibility. If 90% of people cannot reliably distinguish between synthetic and real footage, and if the content generated in certain categories is more convincing than reality, detection is a poor strategy for trust and verification.
conclusion
Assuming we continue to expand training data and computing, video generation models will continue to improve exponentially. The AI industry and society as a whole has reached a tipping point where it is no longer possible for the public to tell whether a video is generated by AI or not.
From photography to Photoshop to traditional CGI, technology has consistently changed public opinion about what makes content “real.” As AI models continue to improve, we expect similar changes to occur again. We believe that fundamental model developers, including Runway, have a responsibility to foster public conversations about the quality of model output and explore ways to push the boundaries of AI research and innovation while mitigating the societal challenges posed by this technology.
All output generated by Runway includes C2PA metadata, allowing you to prove the origin and provenance of the content your model produces. Although this open technology standard has been adopted by various media companies and news organizations, it is not foolproof. We need to build a new, more capable standard that allows for creative possibilities while maintaining trust. It requires not only technological solutions like C2PA, but also new literacies, updated editorial standards, and an ongoing dialogue around trust.
Going forward, we will be working on three principles. That means being transparent about how our models work, collaborating with industry partners on validation standards, and engaging directly with creators, companies, and policymakers to establish new norms for synthetic media.
methodology
Source videos were sampled from Filmpac across five content categories: faces, human full body motion, animals, natural scenes, and urban environments. For each category, we’ve selected representative examples of the content that people often aim to generate. The first frame of each video was extracted and used as input to Gen-4.5 with default settings. Each video is generated once and is not regenerated or post-processed. The actual and generated clips were cropped to 5 seconds and matched in resolution. Participants can watch up to 10 seconds of each video before making a decision. Participants who achieved greater than 75% accuracy (accuracy rate 15/20 or higher, p < 0.05, binomial test) were classified as successful detections.
