The first time I got a chance, I downloaded the Sora app. I uploaded an image of my face (the one my kids kiss at bedtime) and my voice (the one I use to tell my wife I love you) and added it to Sora’s profile. I did this to use Sora’s “cameo” feature to create a silly video of my AI self being shot with paintballs by 100 elderly care facility residents.
What did I just do? The Sora app is powered by Sora 2, an AI model. To be honest, it’s quite a breathtaking model. You can create videos of any quality, from the mundane to the most diabolical. It is a black hole of energy and data, as well as a distributor of highly questionable content. Like many things these days, using Sora feels a bit naughty, even if I don’t know exactly why.
So, if you just generated a Sora video, we have all the bad news for you. By reading this, you want to feel a little dirty and guilty, and your wish is my command.
The amount of electricity used earlier is as follows
According to CNET, a single Sora video uses approximately 90 watt-hours of power. This number is an educated guess derived from research into GPU energy usage by Hugging Face.
OpenAI doesn’t actually publish the numbers needed for this study, and Sora’s energy footprint has to be inferred from similar models. By the way, Sasha Luccioni, one of the hugging face researchers who conducted this study, is not satisfied with these estimates. “We should stop trying to reverse engineer numbers based on hearsay,” she told MIT Technology Review, adding that we need to put pressure on companies like OpenAI to release accurate data.
In any case, different journalists have made different estimates based on Hugginface’s data. The Wall Street Journal, for example, estimates it’s between 20 and 100 watt-hours.
CNET likens that estimate to using a 65-inch TV for 37 minutes. The magazine likens the Sora generation to cooking a steak from raw to rare on an outdoor electric grill (because apparently such a thing exists).
To make your mood even worse, it’s worth clarifying a few things about this energy use issue. First of all, what we just outlined is also known as inference energy consumption. Run the model in response to prompts. The actual training of the Sora model required an unknown but certainly astronomical amount of power. GPT-4 LLM will require an estimated 50 gigawatt hours, which is reportedly enough to power San Francisco for 72 hours. Since Sora is a video model, it took longer than that, but it is unclear how much longer it will take.
From one perspective, you are assuming some unknown cost when choosing to use a model before generating the video.
Second, separating reasoning and training is important in another way when trying to determine how much eco-guilt you should feel (sorry for asking yet?). You can try to abstract away high energy costs as something that has already happened. For example, if the cow that was in your burger died a few weeks ago, or you order a Beyond patty while you’re already seated at the restaurant, you can’t bring back the cow that died. In that sense, running a cloud-based AI model is like ordering waves and turf. All the training data “cows” may already be dead. However, the “lobster” for a particular prompt remains alive until you send the prompt to the “kitchen”, which is the data center where inference is made.
The amount of water used earlier is as follows.
Sorry, I’ll try to guess further. Data centers use large amounts of water for cooling through closed-loop systems or evaporation. There’s no way to know which data center or data centers were involved in creating the video of my friend farting to the song “Camptown Race” as a contestant on American Idol.
But it will still probably contain more water than you feel comfortable with. OpenAI CEO Sam Altman claims that a single text ChatGPT query consumes “about a quarter of a teaspoon,” and CNET estimates that the energy cost of video is 2,000 times that of text generation. So the answer scribbled on the back of an envelope could be 0.17 gallons, or about 22 fluid ounces. This is a little more than a plastic bottle of Coke.
And that’s if you take Altman’s words at face value. It could easily be more than that. Additionally, the same considerations regarding training costs and inference costs that applied to energy usage apply here as well. In other words, using Sora is not a wise choice for water.
There’s a small chance that someone could create a really scary deepfake of you.
Sora’s Cameo privacy settings are robust as long as you are aware of and utilize them. “Who can use it” settings more or less Don’t let your likeness become a public toy unless you select the “Everyone” setting. This means anyone can create your Sora videos.
Even if you don’t want to make your cameo public, the Cameo Settings tab provides additional controls, including the ability to verbally describe how it should appear in your video. You can write anything you like here, such as “I’m toned, toned, and athletic” or “I always pick my nose.” You can also set rules about what you should never show what you’re doing. For example, if you keep kosher, you should never be seen eating bacon.
However, even if you don’t allow others to use Cameo, you can still gain some peace of mind by taking advantage of unlimited features that allow you to create guardrails when creating your own videos.
But Sora’s general content guardrails aren’t perfect. According to OpenAI’s proprietary Sora model card, offensive videos could slip through if someone strongly urges them.
This card displays success rates for different types of content filters ranging from 95% to 98%. However, if we subtract only failures, the likelihood of sexual deepfakes is 1.6%, the likelihood of videos containing violence or gore is 4.9%, the likelihood of what we call “violent political persuasion” is 4.48%, and the likelihood of extremism or hatred is 3.18%. These possibilities were calculated from “thousands of hostile prompts collected through targeted red teaming operations,” in other words, intentionally trying to break guardrails with rule-breaking prompts.
So the chances of someone creating a sexual or violent deepfake of you aren’t high, but OpenAI (perhaps wisely) didn’t say never.
Maybe someone will make a video of them touching poop.
In my testing, Sora’s content filter generally worked as advertised, but I never checked to see what the model card said about its failures. I didn’t painstakingly create 100 different prompts to trick Sora into producing sexual content. If you request a naked cameo of yourself, you’ll receive a “Content Violation” message instead of a video.
but, some Potentially objectionable content is so weakly regulated that it is not completely filtered. Specifically, Sora seems to be indifferent when it comes to scatological content, producing such material without any guardrails unless it violates other content policies, such as policies regarding sexuality or nudity.
That’s right, in my tests, Sora produced cameo videos of people interacting with poop, including scooping it out of a toilet with their bare hands. For obvious reasons, I’m not going to embed the video here as a demonstration, but you can test it yourself. No tricks or quick engineering were required.
In my experience, past AI image generation models, including Dall-E, Bing’s version of OpenAI’s image generation feature, have measures in place to prevent this kind of thing, but that filter appears to be missing in the Sora app. I don’t think it’s necessarily a scandal, but it’s awful!
Gizmodo has reached out to OpenAI for comment on this matter and will update if we hear back.
Your funny video could be someone else’s viral hoax.
Sora 2 unleashed a vast and endless world of misinformation. As an astute and internet-savvy content consumer, you would never believe that something like the viral video below is real. Spontaneous footage appears to have been taken from outside the White House. In what appears to be a wiretapped phone conversation, an AI-generated Donald Trump instructs unknown parties not to release the Epstein files, shouting, “Just keep them out. If I go down, I’ll take all of you with me.”
Judging by the comments on Instagram, some people seem to believe it’s real.
The creator of the viral video never claimed it was real and confirmed to Snopes that it was created by Sola, telling Snopes that the video was “completely generated by AI” and created “solely for artistic experimentation and social commentary.” A likely story. This was clearly created for influence and social media visibility.
But when you publish a video to Sora, other users can download it and do whatever they want with it. This includes posting on other social networks and pretending to be real. OpenAI very consciously made Sora a place where users can doomscroll endlessly. Once you place your content in such a location, context no longer matters and there is no way to control what happens to your content next.
