Luma Labs, the artificial intelligence company that previously launched Genie-generated 3D models, has now entered the world of AI video with Dream Machine, and it's an impressive feat.
Demand to try out the Dream Machine has overloaded Luma's servers, forcing them to put a queue system in place: I waited all night for the prompt to video in, but the actual “dreaming” process takes about two minutes after you reach the front of the queue.
Some of the videos shared on social media by those who were given early access were so impressive that they seemed like they had been handpicked to showcase the best of existing AI video models, but when I tried them out, they were just that good.
It's not at Sora level, and apparently not as good as Kling, but from what I've seen, it's one of the best AI video models to date when it comes to tracking prompts and understanding behavior. In some ways, it's much better than Sora and is available for anyone to use right now.
Each video generation is about five seconds long, nearly double the length of Runway or Pika Labs videos without the extension, and some videos contain multiple shots.
What is your experience using the Dream Machine?
I created a few clips during my testing — one took me about three hours to complete, the other took most of the evening — and while a few of them have issues with blending and blurring, for the most part it captures movement better than any other model I've tried.
I also demonstrated movements like walking, dancing, and running. The old model had people moving backwards or dolly zooming to a stationary dancer from prompts requesting that type of motion. Not so with Dream Machine.
Dream Machine did a good job of capturing the concept of subject movement without needing to specify areas of movement – it's particularly good at execution – but there's very little fine-tuning or fine control beyond the prompts.
This may be because it's a new model, but everything is handled by prompts and the AI uses its own language model to improve automatically.
This is a technique also used by Ideogram and Leonardo to generate images, and it helps to give a better description of what you want to see.
This could also be a feature of a video model built on trans-spreading technology rather than straight-spreading. UK-based AI video startup Haiper has also said its model is best left to prompts, with Sora said to be little more than a simple text prompt with minimal additional controls.
Testing the Dream Machine

I came up with a series of prompts to test Dream Machine, and compared some of these to existing AI video models, none of which achieved the same level of motion accuracy or realistic physics.
In some cases, we gave them simple text prompts to enable the enhancements, in other cases we gave them longer prompts ourselves, and in some cases we gave them images generated by Midjourney.
1. Run for ice cream

For this video, I created a longer form and explanatory prompts – I wanted to create something that looked like it was filmed on a smartphone.
Prompt: “An excited child runs toward an ice cream truck parked on a sunny street. The camera follows closely behind the child, capturing the back of the child's head and shoulders, his excitedly waving arms, and the colorful ice cream truck approaching. The video has a slight jitter to mimic the natural movement of running while holding a cell phone.”
Two videos were made: In the first, the ice cream truck appears to be trying to run over a child, and the child's arm movements are a bit odd.
The second video was much better – it certainly wasn't as realistic and had noticeable motion blur. The video above is of the second shot and also captures the idea of a slight bounce to the camera movement.
2. The appearance of dinosaurs

This time I gave Dream Machine a simple prompt and told it not to expand on the prompt, but to just take what it was given. It actually created two videos that flowed into each other, like the first and second shots of the scene.
Prompt: “A man discovers a magical camera that can bring any photo to life, but chaos ensues when he accidentally takes a photo of a dinosaur.”
Though there is some distortion, especially around the edges, the movement of the dinosaur charging into the room reflects real-world physics in an interesting way.
3. Calling on the Street

Next up is another complex prompt; specifically, one that requires the Dream Machine to take into account light, swinging movements and a relatively complex scene.
Prompt: “A person walks down a crowded city street at dusk, holding a smartphone vertically. The camera captures his hand waving lightly as he walks, showing shop windows, passing people, and the glow of street lamps. The video is handheld and slightly wobbly to mimic the natural movement of holding a smartphone.”
There were two ways to do this: the AI could capture the view from a camera held in a person's hand, or a person walking around holding a camera — first person and third person perspective. The AI chose the third person perspective.
It wasn't perfect, with a bit of crooked edges, but it was better than I expected, considering the mismatched elements of the prompt.
4. Dancing in the Dark

I then started with an image of a silhouetted dancer generated by Midjourney, which I then tried using with Runway, Pika Labs and Stable Video Diffusion, all of which showed movement in the shot but the character was not moving.
Challenge: “Create a mesmerizing tracking shot of a woman dancing in silhouette against a contrasting light background. The camera follows the dancer's fluid movements and keeps her silhouette in focus throughout the entire shot.”
It wasn't perfect – legs would get weirdly distorted when rotating and arms seemed to merge with the fabric – but at least the characters moved, as always with Luma Dream Machine, which offers much better animation.
5. Moon Cat

One of the first prompts I'll try in the new generative AI image or video mode is “Cat dancing on the moon in a space suit,” which is weird enough that I can't reference any existing videos, and complex enough that the video will struggle to follow the movements.
The exact prompt for the Luma Dream Machine was “A cat in a space suit dancing with a dog on the moon.” No fine tuning or explanation of the action types was needed – we just let the AI do it.
What this prompt tells us is that we need to tell the AI how to interpret the movement. It doesn't do a bad job, and it's better than other models currently available, but it's far from perfect.
6. Market visit

Here's another piece that started off as an image from Midjourney – a photo of a bustling European food market. Midjourney's original theme was “highly realistic smartphone photos of bustling outdoor farmers markets in quaint European town squares.”
For the Luma Labs Dream Machine, I simply added the instruction “Walk through a crowded, bustling food market.” There were no other movement commands or character instructions.
I wish I had been more specific about how the characters should move. It captured the camera movement very well, but the figures in the scene were distorted and blended together quite a bit. This was one of my first attempts, so I didn't try any better techniques to move the models.
7. Ending a chess game

Finally, I decided to throw a whole curveball at the Luma Dream Machine. I was experimenting with another new AI model, Leonardo Phoenix, which promises an amazing level of prompt following. So I created a complex AI image prompt.
Phoenix did a good job, but since it was just an image, I decided to put the exact same prompt in Dream Machine: “A surreal weathered chessboard floating in misty space, adorned with brass cogs and gears and populated with elaborate steampunk chess pieces, including steam-powered robotic pawns.”
Ignoring almost everything else but the chessboard, he created a surreal video of the chess pieces being blown off the edge of the board as if they were melting. I'm not sure if this was intentional or just a failure to understand the movements, as there is an element of surrealism to it. But it looks cool.
Final thoughts
I did some math and found that I accessed the Luma Dream Machine on a Saturday evening, used it for a few days, and made 633 generations. Of these 633 generations, I'd say at least 150 were just random tests for fun. So I estimate it took about 500 generations… https://t.co/TpMCdDmlxyJune 12, 2024
Luma Labs Dream Machine is an exciting next step in generative AI video. It appears to have leveraged its experience in generative 3D modeling to improve motion understanding in video, but it still feels like a stepping stone to true AI video.
Over the past two years, AI-generated images have evolved from strange, low-resolution renditions of multi-fingered, multi-faced humans that look more like Edvard Munch paintings than photographs to something nearly indistinguishable from reality.
AI video is much more complex: it doesn't just have to recreate photorealism; it has to understand real-world physics and how it affects scenes and the movement of people, animals, vehicles and objects.
For now, even the best AI video tools are meant to be used alongside traditional filmmaking, not replacing it. But we’re on the brink of Ashton Kutcher’s prediction of a time when everyone can make their own feature films.
Luma Labs has created one of the most realistic motion tools I've seen, but it's still not at the level I need it to be. I don't think it's at Sora level, but I can't compare it to the videos I've made myself with Sora. I've only seen videos by filmmakers and OpenAI themselves, which were probably culled from hundreds of failed attempts.
Abel Art, an avid AI artist who had early access to the Dream Machine, created some impressive work, but he said he had to create hundreds of generations of even a one-minute video to ensure consistency and discard unusable clips.
His ratio is roughly 500 clips per minute of video, with each clip lasting about five seconds, and he discards 98% of shots to create the perfect scene.
The ratios for Pika Labs and Runway are likely even higher, but reports, at least from filmmakers who have used Sora, suggest that Sora's waste rate is comparable.
For now, even the best AI video tools are meant to be used alongside traditional filmmaking, not replacing it. But we’re on the brink of Ashton Kutcher’s prediction of a time when everyone can make their own feature films.
