Last week, Google began rolling out its VEO 2 video generation model to Gemini Advanced subscribers. I've been playing with it ever since. That's very true, I already seem to be against Google's monthly restrictions for the video generation.
With the broader stages of Veo 2 availability, Google naturally highlighted clips produced by models that are difficult to distinguish from human videos, whether they are asked to mimic realistic footage or cute animations. The results I saw in my time with Veo 2 were less impressive, but I have to say it's closer than I expected. Here are five of my favorite early results for VEO 2:
5
Shark party in the forest
In this video I engaged in the old tradition of inviting generative AI to create meaningless flow of consciousness style scenes. We asked Veo 2 for a video of a hybrid of embarrassment of a human being holding a red cup with a bonfire in the woods. Please check and check. Also, that prompted me to include a “van” in the scene. Assuming you know from the context, it means the type of van that someone might go to camp. Instead, we gave all the sharks recognizable bun sneakers. It's not what I wanted, but it's definitely fun.
The 8-second clip of a Sharkman dancing around the fire becomes realistic at a glance. The fire burns realistically, the background is compellingly blurry, and the shark skin shows a realistic texture. The details are not very clean. The shark appears to have a regular flipper and one humanoid hand. The cups of the cups in the background are floating near the hands rather than being held within them. Still, a stunning effort for something very meaningless.
4
Etched golden skull
To feel how the VEO 2 manages complex textures, I rotated a gold skull with finely etched details in the style of a calavera, under bright light. The results feel imperfect, with the skull partially turning, pausing and continuing, but both the anatomy of the skull and the way it plays light in its various textures and details is compelling.
3
Gen-Ai Newscasters
Thinking about how video generation could potentially be used to deceive people, I urged VEO 2 to simulate cable news broadcasts, and the anchor sat at the desk and spoke to the camera. In most cases, the results are convincing. One anchor speaks as another nod. They even have realistic reflections on the surface of the desk.
But Veo 2 fell into this one text. Chyron at the bottom closes “AI video generation here. What is that for?” but it's not perfect. The small details of the video are also slightly off. The other wears two microphones so that one anchor pen appears and disappears. The graphics behind the Anchor feel appropriately cheesy in the TV news segment on AI video, but they feature film reels overlaid with many things and zeros.
2
The Legend of Zelda – Aaaah?
I was curious if Veo 2 was trained with video game footage. I've encouraged to create a scene from some specific titles to find it. Now I've explained to you about the opening moment of The Legend of Zelda: Breath of the Wild. There, Link flows out of the cave and gazes at the scenery from the cliff.
VEO 2 couldn't do that specifically, but it's absolutely trained with the video of the game. Characters like Link are run off into caves and cliffs. There are plenty of items on the back, and if you squint, it looks vaguely like a sword and a shield. Interestingly, the game's UI is almost unharmed. All the elements are in the right place, and the corner maps will realistically rotate when the camera pans.
1
Clear Cyberpunk 2077 Gameplay
After Breath of the Wild, I urged VEO 2 to some more game footage, but the simple prompt, “Cyberpunk 2077 Gameplay,” produced what seemed like the most accurate result. Rainy City Street, UI, a small aircraft – it's all very similar to cyberpunk. There are even signs promoting what looks like a cybernetic implant.
The more detailed details are confusing. The text and iconography were wavy and vague, and Veo 2 appeared to be throwing an animation that throws a gun, even though the player's characters weren't actually moving through the scene. Still, Veo 2 knows what Cyberpunk 2077 looks like and isn't afraid to replicate it.
Affordable AI video generation is here
My first week with Gemini's Veo 2 was very similar to my early experience with the AI Image Generation app. The novelty of plugging in at a short prompt means getting a short video in 1-2 minutes. This means that even if the results are less than the stars, it is still interesting. It's new, weird and fun.
But I don't know exactly what to do with VEO 2 other than around Goof. Given how resource-intensive AI video generation is, offering VEO 2 as part of a $20/month subscription seems to be completely unsustainable for Google. Additionally, Gemini may ultimately offer “freemium” video generation capabilities that don't cost anything at all. Gemini Advanved cut me off after generating about 50 clips. Potential free versions of the feature are further limited.
Whatever Google's long-term video generation ambitions, Veo 2 is currently widely deployed for Gemini Advanced Subscribers in both mobile apps and Gemini web interfaces.
