I tested SORA 2 against Google's VEO 3, and the gap is incredible

AI Video & Visuals


Google Pixel 10 Pro XL Gemini Home

Ryan Haynes/Android Station

If you purchase a Pixel 10 Pro series phone or last year's Pixel 9 Pro, you can get a year's worth of Google's Gemini Pro subscription. This $20/month service unlocks a powerful Gemini 2.5 Pro model and a suite of cutting-edge AI tools. Until recently, the crown jewel in this package was the Veo 3, Google's impressive text-to-video generator.

However, the world of AI moves at the speed of lightning. Over the past week, Openai has unveiled two competing Sora models. This means that Google's video generator is no longer the only game in town. SORA 2 is only invites for now, but the model already has an active user base. Naturally, I took Openai's Sora 2 on Spin vs Google's VEO 3 to see which AI video generators dominate.

Google Veo 3 vs Openai Sora: The results are amazing

Start with a simple prompt with no characters or complicated details that can trip any of the AI ​​video generators. Given the static nature of this shot, you would expect every model to nail the task. However, the results were significantly different.

The attempts of the first generation Sora model were visible at a glance. Understanded objects such as cups, liquids, machines, and other things and assembled them in the correct order. However, the fantasies quickly fell apart. The “espresso” was thick, blurry and consistent, splashing into the cups with unnatural physics. It was a video of the words of the prompt, but there was no sense of artistry or realism.

In contrast, the Veo 3 generation felt like they were captured by professional videographers. The espresso flowed with a compelling viscosity, and the liquid swirled realistically as it settled. The coffee is dispensed only from one side of the portafilter, so it's not a perfect result, but it's a huge improvement on Sora's attempts.

The SORA 2 is the latest and greatest bundle. We present realistic physics without the errors shown in the Veo3 results. But is that a massive improvement? not much. But fortunately, for Openai, it's just starting out.

How about animals? The first generation SORA model did an acceptable job of capturing the enthusiasm of golden retrievers in actual busy parks. VEO 3 did a bit of a good job, but the random sea of ​​background characters was a clear indication of the existence of AI.

Sora 2 is a place where things get you worried. The golden retriever was rendered extremely accurately and the entire scene was trusted. People in the park were not blurred. My only knit pick is that there are too many other dogs in a normal city park for the scene.

We then asked for motorcyclists riding along the beach at sunset. Again, the original SORA model gave us a cartoonish result of boundary lines where one motorcycle fishes while another motorcycle glides into the water with zero resistance. I cannot inherit this outcome. Surprisingly, SORA 2 unexpectedly failed on this task, making the same mistake as its predecessor.

Meanwhile, Veo 3 delivered shots that looked completely film-like. The motorcycle moved on the sand as expected, leaving behind tread marks and dust marks, and the bike was slightly tilted as the riders turned. But the lighting was the most amazing part. The low sun casts long dramatic shadows and shines realistically from the motorcycle.

My next prompt proved to be a difficult challenge for older models. Sora and Veo 3 were unable to generate clips that are available, but their obstacles were still interesting.

Sora's version broke the rules of reality. It struggled with the persistence of the object, and pedestrians were present on the sidewalk or temporarily fused with each other at a jarring moment. Needless to say, this dreamy sequence doesn't resemble reality.

The VEO 3 attempt was more consistent, but failed to execute the details. It did a much better job capturing the authentic vibe of Kolkata, but the taxi itself moved in a strange slide motion that didn't feel a connection to the road. Furthermore, all text was rendered unreadably, as is common in AI. The new SORA 2 models performed much better, nailing the urban atmosphere and even the vehicle residents. It can be easily handed over as an actual video.

Finally, let's take a look at what I think is the most impressive result of Google's model, Bangkok's Mandalorian. Surprisingly, neither Sora nor Veo 3 rejected my prompt on copyright grounds.

Either way, the VEO 3 results were incredible. The characters it created were real deal split images, from the particular sheen on the armor to the iconic silhouette of the helmet. It didn't look like the AI ​​generation, it looked like a scene removed from the show.

Sora, on the other hand, provided at best a close approximation. It produced a common character covered in shiny polished chrome with neon lights reflecting from the surface. It captured the Bangkok part of the prompt, but failed on the main subject. In a way, Sora avoided copyright infringement, but it also failed to follow my instructions exactly.

Unfortunately, the new Sora 2 model refuses to generate videos containing copyrighted characters.

AI video generation has come a long way

Ask Google Gemini to generate videos using Veo 3

Mishaal Rahman / Android Authority

When Openai unveiled SORA in early 2024, most of us were surprised at how realistic and persuasive it was. These early samples showed off impressive film talent and promised to disrupt video production. At the time, Openai also had one of the best AI image generators in the form of Dall-e. However, when Sora finally launched in December 2024, it didn't reach those lofty expectations. Nonetheless, Google followed up the VEO model just a few days later, steadily repeating the aggressive updates that culminated in today's VEO 3.

Unfortunately, Google's early release of AI video generators was not as perfect as the demo suggested. However, Veo 3 and Sora 2 are completely different beasts.

Early VEO and SORA models suffered from the same indications of the generator AI. Background objects shift unnaturally, characters do not have object persistence, sometimes blending into the environment or blending with each other. Physics was barely important and fortunate to have the consistency of the story, as objects moved in a frictionless, impossible way.

SORA 2, and Google's VEO 3, address almost all of these flaws. A single sentence prompt can now generate authentic videos. This makes these AI video generation tools extremely useful for creating light content. Teachers can create visual stories for the class, and business owners spin quick ads on social media. The use cases seem endless.

The only problem is cost. With Gemini Pro, you only have 3 VEO 3 videos per day. However, it turns out that a Google Labs project called Flow also grants 1,000 AI credits per month. This is converted to around 100 videos using the “fast” model of VEO 3.

Meanwhile, SORA 2 is currently free to use without a ChatGPT subscription. Openai CEO Sam Altman acknowledges that this open access is unsustainable. This is because use has already exceeded expectations. Daily restrictions seem inevitable, but to be fair, I got a clip that can be used on my first attempt, thanks to the model's stronger grasp of physics, movement and the nuances of real world.

The catch is that Sora 2 hasn't been released yet, and Openai almost certainly places strict limits on the number of video generations when the service is deployed more widely. So for now, VEO 3 is one of the most protected secrets of Google's Gemini Pro subscription.

Thank you for being part of our community. Please read our comment policy before posting.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *