What exactly does “video” mean when AI can make movies?

AI Video & Visuals


For the past few weeks, I've been using Apple's iMovie software to create home videos on my phone. The idea is to weave together clips of the family shot in February. We plan to continue working on this until March. So far, the film has shown her 5-month-old daughter cooing and waving her arms. My 5-year-old son was chasing me with a snowball. And above all, a visit to our town's creepy dilapidated amusement park.

Yesterday, while listening to the announcement of Sora, an amazing new text-to-video conversion system from ChatGPT developer OpenAI, I was reminded of my own movie. Sora can take prompts from users and create detailed, original, and photorealistic one-minute videos of him. OpenAI's announcement features many fantastical images, including an astronaut seemingly stranded on a wintry planet, two pirate ships dueling over coffee, and “historic footage of California during the gold rush.” A video clip was introduced. However, his other two clips were more intimate, the kind that could be captured on an iPhone. The first was generated by a prompt asking for “a beautiful homemade video of the people of Lagos, Nigeria in 2056.” If that's true, it “captures” what appears to be a group of friends or relatives sitting at an outdoor restaurant table. The camera pans from a nearby open-air market to a cityscape divided by a busy highway at dusk. The second photo is “Reflection in the window of a train running in the suburbs of Tokyo.'' It looks like the footage any of us would have taken on a train. The silhouettes of passengers can sometimes be seen superimposed on the passing buildings on the window glass. Strangely, no one seems to be taking pictures.

These videos are flawed. Many are too perfect and have a slightly cartoonish quality. However, there also seem to be some that capture the texture of real life. The magic behind this is too complex to explain easily. Broadly speaking, it may be correct to say that Sora does with video what his ChatGPT does with writing. OpenAI claims that Sora “understands not only what the user asks for in a prompt, but also how those things exist in the physical world.” It figures out how different kinds of objects move and interact with each other in space and time in a statistical, mind-adjacent, and (possibly) unconscious way. Sola “might not understand specific cases of cause and effect,” the developers wrote — “For example, if a person bites into a cookie, the cookie may not leave a bite mark afterwards.” Maybe not.” Still, the comprehensive understanding of the objects and spaces that AI invokes means it's more than just a system that generates video. This is a step “toward building a general-purpose simulator of the physical world.” Sora performs his work not only by manipulating pixels, but also by conceptualizing his three-dimensional scenes as they unfold over time. Our own minds probably do something similar. When we envision landscapes and places in our mind's eye, we are not only imagining what they look like, but what they are.

Currently, the system is available for testing to a small number of professionals, but not to consumers. OpenAI is announcing this as a preview “to give the public an idea of ​​what AI capabilities are coming.” I watched the demo video and wondered what I could do with it myself. Could we ask a future version of Sora to generate a clip for the Home Girl movie? “A 5-month-old girl in a red sweater waving her arms and imitating her brother saying “Lego'' Could you please show me the phone video of you saying that? What if an AI could access my previous home movies and photo library showing my home and family from different angles? Would I want to allow that access? AI Source Material It's a little strange to think that is not just a photo, but an idea: an idea of ​​Lagos or Tokyo, an idea of ​​a family or group of friends, an idea of ​​a “beautiful handmade video.” Sora is not Photoshop. Sora contains knowledge about what Photoshop shows us.

How might “synthetic” videos of the kind produced by Sora and his descendants be used? Perhaps bad actors could create deepfakes to spread misinformation or use them as anti-personnel weapons? There is a gender. Businessman inserts composite clip into presentation. Filmmakers, advertisers, and studio executives can use composite video to storyboard their ideas and even, if film industry unions permit, use composite video to create complete shows. You can also create one. New and unfamiliar creative enterprises will arise, which are difficult to imagine today. It's a path to entertainment, education, and diversion that we can't yet imagine. (If we knew what those were, we'd all be millionaires in the future!) Assuming lawsuits don't stop these systems from incorporating copyrighted material; Visual styles invented by famous artists are called upon by prompts and can become boring if used too much. Shots that are currently expensive and time-consuming to produce could become cheap and instantly available. There is no problem even if the script starts with “EXT.”. A market in Lagos. ”

Inevitably, the meaning of video as a medium will also change. Perhaps we will begin to suspect that all videos are synthetic and will stop trusting them. We may choose not to distinguish between what is synthetic and what is not true. In 2018, Peter Jackson released a documentary about World War I called They Shall Not Grow Old, which consisted entirely of colorized archival footage. Since reality is in color, the retouched footage was probably closer to reality than the black and white film it was based on. “We tried to keep it real,” said John Newell, one of the film's colorists. The AI ​​also tries to keep it real. If the synthetic video is based on statistical extrapolation of the audio, you may end up thinking the synthetic video is real enough.

Photographing things may start to feel superfluous. On YouTube, the channel Beirut Explosion Angles collects footage of the massive explosion that devastated the city of Beirut on August 4, 2020. There are currently 938 clips of explosions, many of which were captured with cell phone cameras. The clips often feature people holding their phones aloft to film the explosions, allowing for more angles to be collected. But if you want to see what an explosion looked like coming from a particular address, composite footage is available. It is based on a large amount of data and is realistic enough. When OpenAI researchers asked Sora to create “an aerial view of Santorini at blue hour, showcasing the stunning architecture of the white Cycladic buildings with blue domes,” tourists The resulting image was almost indistinguishable from the one taken. So why launch a drone in the first place? AI “understands” the concept of Santorini, and so do we.

Yesterday morning, my son was giving my sister a strange look. She smiled at him from above her bodyguard. I reached for my phone to capture the scene, but then remembered that I had put my phone in another room to avoid distractions. I know this moment happened. You can picture it in your head. It would be a great clip for my home movies. There is a sense in which we can “prompt” our brain (“Remember when?”) and generate memories accordingly. So why not prompt the AI? What's wrong with a real fake video?

Probably nothing. However, a composite video is not a recording, so it is different from all the clips you have previously captured. It becomes a rendering of an idea.

What makes AI so powerful today is ideas. Artificial intelligence relies on the fact that everything is information. The position of the pieces on a chessboard, the style of an author you admire, the look of Lagos at dusk can all be explained through text, images, video and audio, but to a large extent they are medium agnostic. Those are ideas. When you write down an idea, draw a picture, or photograph something, it may seem static and solid, but it's actually fluid. There's always another way to write. That photo could have been taken from a different angle. Books can also be made into movies. The same prompt can be slightly modified to pull short stories from ChatGPT, comic panels from DALL-E, video clips from Sora, and so on. When prompting the AI, there is no need to go into details. AI more or less understands what you want to say.

In the 1920s, Karl Ove Knausgaard's six-volume work Mein Kampf became a literary sensation. Although the work was marketed as fiction, many of the events depicted actually happened, and many of Knausgaard's relatives have been identified by their real names. Was it a novel or a memoir? Is it fiction or non-fiction? It was a little bit of all those things. We intuitively understand that text is always an expression of ideas, and that ideas are fluid. We know that books are not records, and that writing is always slippery.

We tend not to have textual intuitions about other media, mainly for logistical reasons. Forging a video has always been more difficult than editing a document. But we need to develop it. What's true for text is also true for audio, video, and any other kind of expressive artifact. If you want to know if a book is true, you have to look outside the book. I can't believe how bookish it is. On the other hand, books move us, sometimes happily, beyond expression and into imagination. Apparently everything is heading there. ♦



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *