AI video platform Runway plans to release its third-generation model “in the coming days,” which will offer “significant improvements in fidelity, consistency and motion compared to previous-generation models,” as well as being significantly faster, the company told Tom's Guide.
Runway released Gen-2, the first commercial AI model that generates video from text, in June of last year. Since then, it has unleashed a revolution in synthetic video onto the world. It is now competing with companies like Pika Labs, Haiper, Luma Labs, and the yet-to-be-released Sora.
Gen-3 marks a major shift for Runway and the AI video space. It has been rebuilt from the ground up using a new generation of infrastructure purpose-built for large-scale multi-modal training. The new models are trained on images and videos simultaneously for improved realism.
The public will have access to the alpha version “in the coming days,” and Anastasis Germanidis, CTO and co-founder of Runway, said it's the smallest of a new generation of cutting-edge AI models that will emerge as a result of the new training infrastructure.
What's different about Runway Gen-3?

Runway Gen-3 has improved capabilities for controlling motion in video and understanding real-world movement and physics, which when combined with photorealism results in a model that allows you to create videos that are nearly indistinguishable from reality.
Gen-3 Alpha shows significant improvements in terms of temporal consistency and significantly reduced morphing compared to Gen-2 for both text and image inputs.
Anastasis Germanidis, CTO of Runway
After completing training, the team was surprised by several things when using Gen-3 for the first time, including its approach to scene creation, which was made possible by the ability to create a minimum of 10 seconds of video, whereas the previous generation capped it at around 4 seconds.
“The ability to create unusual transitions has been one of the most fun and surprising ways to use Gen-3 Alpha internally,” said Germanidis. “The model allows us to incorporate and understand dramatic changes in the environment, with very satisfying results.”
Introducing Gen-3 Alpha: Runway's new base model for video generation. Gen-3 Alpha is capable of producing highly detailed videos with complex scene changes, a wide range of cinematic options, and detailed art direction. https://t.co/YQNE3eqoWf(1/10) pic.twitter.com/VjEG2ocLZ8June 17, 2024
In addition to varying scenes and environments, “because it is trained on multiple highly descriptive captions per scene, it can generate videos with unusual and interesting transitions of environment and action, as well as temporal keyframing of specific elements,” he explained.
“These model improvements, combined with existing control modes like Motion Brush, Advanced Camera Controls, and Director Mode, give users more control than ever before.”
Gen-3 lets you start with images, text, and even video, whereas Gen-2 doesn't support video as an input. Germanidis says it doesn't matter which you use: “Gen-3 Alpha shows significant improvements in terms of temporal coherence, and much less morphing compared to Gen-2, for both text and image inputs.”
Creating a general world model

Gen-3 Alpha by @runwayml is awesome, but it's not a Gen-3 world model without audio 😉🎶Enjoy our updated Runway demo with accurate music and SFX!🚂💨A woman's subtle reflection in the window of a train moving at lightning speed through a Japanese city. https://t.co/Iq293vT7N6 pic.twitter.com/6nOIeEjRAqJune 17, 2024
Germanidis told Tom's Guide that this is “the first of a next generation of foundational models that Runway will train from scratch,” adding that future versions will “reach and exceed the scale of large-scale language models like Google Gemini and Anthropic's Claude.”
The models can struggle with complex character and object interactions, and generation doesn't always follow the laws of physics exactly.
Anastasis Germanidis, CTO of Runway
Just as large AI LLM labs like OpenAI and Anthropic are working towards achieving artificial general intelligence (AGI), Runway is working to build a “general world model.”
“A general world model is an AI system that builds an internal representation of an environment and uses it to simulate future events within that environment,” Germanidis explains.
“The purpose of the General World Model is to represent and simulate a wide range of situations and interactions that would be encountered in the real world,” he added.
Gen-3 isn't an open-world model per se, but it's a first step, Germanidis said. “It's still very early days, and this is the first and smallest of the models we'll be releasing.”
“The model can struggle with complex character and object interactions, and the generation doesn't always follow the laws of physics exactly,” he warned. So don't get too excited and remember that this is just a first step.
