Last year, I deepfaked a kid’s stuffed animal to make it look like the stuffed deer was on vacation.
This was an experiment Google was running to see if they could recreate the events depicted in the Gemini ad, and they never showed a video of Buddy the Deer’s adventures to a 4-year-old. But it was an exercise in discovery that made me think a lot about the difference between harmless fun with generative AI and full-blown slop. Maybe that Venn diagram is a perfect circle. Probably not. But what I do know for sure is that the tools to create realistic videos are amazingly good and require surprisingly little effort and know-how. And this trend will continue hotly until the Omni era of Gemini.
Omni is a new family of generative models that is said to one day be able to transform any kind of input, including photos, videos, and text, into other things. But first, just create a video. Omni Flash is the first of these models released by Google and is now available in Flow, the company’s AI video generation and editing platform. You can still use previous models of Veo if you like, but Omni has improved on Veo in several ways.
Omni allows you to upload a video and use it, along with text prompts, as a starting point for your AI-generated creations. Google also claims that Omni incorporates more real-world knowledge when creating videos, resulting in better character consistency across videos. There’s only one way to really know if these claims are true. I brought my AI buddy back to pack a small AI-generated bag for another adventure.
The results are very complex and perplexing. Some were very good, much more consistent and true to my prompts than when I was testing Veo 5 months ago. But even with the best clips Omni created for me, certain AI jump scares still remain, such as when Buddy suddenly turns around while skydiving.
In another video, I gave Omni some artistic freedom. “We create a montage of Buddy packing up for vacation and boarding a cruise ship for a tropical getaway. It’s a cute, playful vibe. Buddy packs something interesting in his suitcase, which comes later in the clip.” It made Buddy pack a jar of honey. Later in the clip, he reaches for it as if it were a bottle of sunscreen. “Hmm,” the character says, squirting honey on his hoof.
Honestly, it’s not bad. However, the bottle of honey constantly changes throughout the video, going from a jar to a clear squirt bottle with water and back to a squeeze bottle with honey. And I can’t even begin to explain how the model came up with the final frame of the video. It’s as if you barfed up a bunch of elements from the sequence you just created.
Suggest edits to your video using text-based prompts. Kudos to Google. This works better on the Omni than it did when I tested the Veo 3. But the result is: bad The Veo case was so disappointing that I found it much easier to just prompt me to create a new video from scratch every time I wanted to change something. Omni will actually reflect your edits, but the results may not always be a hit.
We had a vacation clip highlight Buddy’s facial reactions, and the results looked weird. Also, Buddy sometimes grows horns, but he doesn’t have them. Buddy is babythank you very much. When I prompted it to remove the horns that appeared in one scene, it complied and then added the horns to all other scenes as well.
The problem is, none of these are free. Generating a video costs credits, which can range from 15 to 40 credits depending on the length of the scene and the “stuff” you start with. Each edit costs 40 credits. I’m on the $20/month AI Pro plan, which comes with 1,000 credits per month. After making some edits and generating about 20 clips, I was down to 145. If you have a specific idea of the video you want Omni to produce, you may find yourself spending a lot of time interacting with the model to get a video that closely matches your vision.
I can honestly say that I was not prepared for what I saw.
One of Omni’s supposed strengths is adding AI-generated stuff to real videos, so I gave Buddy a break and deepfaked himself. I started with a blank selfie and asked Omni to generate videos of me eating spaghetti, sitting in an airplane seat, standing in front of the Eiffel Tower and biting into a baguette. And I can truly say that I was not prepared for what I saw.
My deepfake video has AI talking. The sound of a fork hitting a bowl of pasta is a little over-produced. A woman appears twice in the background of the airplane video. But aside from those minor glitches and a vague sense of eeriness there, they’re pretty convincing.
I showed my husband the pasta clip. He knew I was testing an AI video tool, but I didn’t tell him what was being generated by the AI in the scene. Without knowing what the AI had produced, he assumed I was sitting in front of a camera eating pasta, and said the only clue that something had happened was that the bowl looked strange. Eating pasta itself seemed convincingly real. my husband. Basically men who have seen me in real life every day for the past 10 years.
My other deepfakes are of varying degrees of “good enough to fool people on social media.” Some of the Eiffel Tower clips look a bit cartoonish, but one is so convincing that you might have to watch it a few times to confirm it’s AI. I When the AI “me” turns around and shows me her hair pulled back in a ponytail, I know it’s not me. But I’m not sure if other people can understand the difference and that makes me feel weird.
We are definitely deep in the uncanny valley
To be honest, I’m a little tired of everything. When I tested the Veo 3, I was struck by the realism it produced. Time and time again over the last few years, I have been shocked by how easy it is to create a fake persona with fake photos. I should probably be shocked by Omni too, and I probably am, but the edge has faded.
Creating an AI-generated cinematic masterpiece isn’t as easy as Google hopes. But Omni improves on Veo in some recognizable ways. With just a little effort, if you have a Google account and a credit card, you can take a video of yourself sitting at home and make it look like you’re on a plane to Maui. I don’t think we’re exactly at the “base of the singularity,” but we’re definitely deep in the uncanny valley.
All images and videos in this story were generated by Google Gemini.
