Artificial intelligence models can generate lifelike video footage with simple text prompts. However, these tools still struggle to produce realistic videos of complex natural movements such as human dance.
When CalMatters and The Markup asked dancers and choreographers whether AI could disrupt their industry, most concluded that human dancers cannot be replaced.
It turns out they were right most of the time. We tested nine different cultural, contemporary, and popular dance styles using four commercially available generative AI video models, generating a total of 36 videos. The latest commercially available AI video generation models have produced convincingly lifelike videos of people dancing, but none of them have produced people performing the dances they are instructed to do.
In about one-third of the videos generated, the subjects appeared inconsistently from frame to frame, with abnormalities in their movements and limbs. The frequency and magnitude of observed issues has improved significantly compared to the first test in late 2024.
CalMatters and The Markup tested four commercial video generation models created by leading technology companies to create video clips of traditional and popular dances.
We limited our testing to consumer-grade, closed-source generative video tools. This is because these tools are the easiest to use for everyday users and tend to perform better than open source models. We tested OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
We created nine video prompts that test different dances in a variety of environments, including dance floors, stages, bedrooms, studios, cultural events, public squares, and classrooms. We tested popular contemporary and traditional cultural dance styles, including the Macarena, Mashed Potato, Folklorico, and popular TikTok dances. Please refer to appendix For more information.
We varied the level of specificity to test whether identifying dances by name is sufficient to generate a video of the desired motion, or whether explicitly specifying the exact body movements would improve the output.
Before finalizing the list of prompts, I submitted them to ChatGPT for editing based on the Sora 2 prompt guide. look Limitations: Rapid optimization For more information.
Each prompt was sent once using each model’s default settings to produce a landscape video. The three prompts sent to Sora 2 were edited to remove words that would trigger OpenAI’s filters, blocking prompts that may have violated “guardrails regarding similarity to third-party content.” For example, Sora 2 flagged prompts that mentioned specific years, popular music artists, or banned words. One of the blocked prompts was a video of a politician dancing the Macarena. For that prompt, replacing “politician in a suit” with “man in a suit” got around the guardrail. In Veo 3.1, similar prompts were flagged when sent via Gemini or Flow, but not when sent directly to the Veo 3.1 API.
We evaluated the generated videos based on six different criteria regarding prompt alignment and video consistency.
- Did the main character dance?
- Did the protagonist perform a specific dance that we requested?
- Did the main subject maintain the same appearance throughout the video?
- Did the main subject matter create realistic movements based on human physiology?
- Did the scene and setting match the prompt?
- Did the camera match the indicated camera angle and position?
Each of the above criteria was rated as pass or fail by one reviewer, with assistance from a second reviewer if necessary. The generated cultural dance videos were reviewed for accuracy by dancers familiar with the dance.
Of the 36 videos generated, all but one featured dancing. One video produced by Kling 2.5 did not show her dancing, but instead showed her lower half doing a side lunge.
There was no video produced of the actual dance we requested. Regarding the Cahuilla Band of Indians’ bird dances, tribe member Emily Clark said, “None of these depictions come close to bird dances, in my opinion.” Although the Horton Dance video did not show the specific dance moves we requested, choreographer Emma Andre said she felt the Veo 3.1 depiction was “surprisingly lifelike.”
For the remaining pop culture dances, we compared the generated videos to videos found on YouTube to assess whether the dances were accurate.
Eleven of the 36 videos showed inconsistent movement or appearance. This includes sudden changes in the structure of clothing, hair, and limbs, such as the head rotating on an axis separate from the body or limbs liquefying and reconfiguring.
Please refer to appendix See full results and video.
No images were used to provide instructions to the model. Image-to-video generation involves uploading a static image along with a text prompt and generating a dynamic video from both. Image-to-video generation is an advertised use case for models that generate dance videos from user-submitted images.
We did not request videos featuring multiple dancers, even though some dances are often performed in groups. To avoid ambiguity as to whether the evaluation failure was due to problems in generating complex human movements or realistic multi-subject videos, we restricted the video prompt to showcasing a single dancer.
I haven’t optimized the prompts for each model. Each company publishes its own instant guide. (See guidelines for Veo 3.1, Hailou 2.5, Kling 2.3, and Sora 2.) Instead, we used ChatGPT 5 to standardize prompts across models to align with the Sora 2 prompt guide. Optimizing the model’s prompts by following specific guides could have yielded more accurate results.
We’ve also worked hard to improve the quality of our videos by providing detailed step-by-step instructions for each dance. However, these instructions did not produce more accurate videos than those produced with simpler prompts.
We did not test generative models that focused on human movement generation. These models are used to generate and capture natural human movement in animation and video games. Researchers are using large datasets, including popular dance videos on TikTok, to train the most advanced academic models in the field. These models may perform better than the consumer models we tested, but they require technical expertise and significant computational resources to run.
Our evaluation is limited to videos generated for nine prompts. It is not a comprehensive evaluation of the models used. Some video generation benchmarks, such as Tencent’s AI Lab, use hundreds of prompts to test features such as complex motion, multiple subjects, and creative styles.
We would like to thank Yuhang Yang (University of Science and Technology of China) and Xiaodong Cun (University of Great Bay) for reviewing early drafts of this methodology.
view Rating with prompts or Evaluation by model.
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
prompt: “In a brightly lit dance studio, a woman performs the Summer 2024 ‘Apple’ dance (Charli I removed this reference because it was rejected. (See reference video)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “A Cahuilla Indian woman wearing a colorful ribbon skirt performs the Bird Dance in slow, majestic movements. The camera remains still. The atmosphere is respectful, ceremonial, and visually rich.” (Watch informative video)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “In a well-lit school gymnasium, a teacher in civilian clothes does the chicken dance. It’s fun and silly, with arms flapping and hips twisting. The camera doesn’t move, you just watch them do the moves with their whole bodies.” (See video for reference)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “In a bright, white-walled dance studio, a dancer wearing tights performs Horton Enhancement Number 3 of Lester Horton’s technique: strong lines, extended limbs, and controlled movements. The camera remains stationary, observing the shape and transformation of the body.” (Watch reference video)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “Somewhere in the state of Jalisco, Mexico, in a public plaza illuminated by a golden sun, a woman wearing traditional Jalisco folklorico costume performs a Jalisco folklorico dance. The camera remains still, watching her whole body move from a comfortable distance. The natural light and the light colors of the dress give the atmosphere a festive, free-spirited feel.” (See video for reference)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “On a well-lit dance floor, a stylish politician in a suit dances the Macarena. The camera remains still, capturing his movements in the suit, the retro fun of the moment, and the light illuminating the floor.” In Sora 2, “politician in a suit” is replaced by “man in a suit.” This is because the original prompt was rejected for potentially violating the platform’s guardrails regarding third-party likeness. (See reference video)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “On a brightly lit stage, a man in a shiny satin shirt and flared pants performs the classic 1962 mashed potato dance. The camera stands still and watches as his legs swivel and his body sways in vintage style. The atmosphere is retro, fun, and energetic.” Sora 2 removed references to that year after being rejected for potentially violating the platform’s guardrails regarding third-party similarities. (See reference video)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “In a cozy, blue-walled bedroom, someone in fun pajamas is doing the Renegade dance (a dance made famous on TikTok in 2019). The camera doesn’t move, just watching them hit the moves. The atmosphere is playful and relaxed.” (Watch helpful video)
Clockwise from top left: Video generated using OpenAI’s Sora 2, Google’s Veo 3.1, Kuaishou’s Kling 2.5, and MiniMax’s Hailou 2.3.
markup
prompt: “On a dark, spotlight-lit stage, a figure in a sharp business suit performs the classic sprinkler motion: one hand behind his neck, the other spinning lightly in a wide arc. The camera remains still, capturing the whole body. The atmosphere is retro, playful, and fluid.” (See video reference)
