Before we begin, let’s take a look at an animated GIF.
(Image source: Bilibili)
That’s a really cool movie scene, right? The material and atmosphere are perfect.
But what if I told you that this movie was entirely generated by AI? I think many readers might be surprised or even try to go back and find any flaws in the clip.
In recent years, with the rapid development of technology, it has become increasingly difficult to distinguish between special effects and AI. It seems like creating your favorite videos has never been this easy.
However, like me, I think most people just watch and don’t practice, or have tried but given up before getting serious.
The reason, in a nutshell, is as follows. This kind of thing is really depressing.
If you want a higher level of completeness, you should deploy your own models and set up stable and controllable workflows on top of ComfyUI. However, these precise parameters are still a mystery even to me, who has been involved in the AIGC field for many years. It is no exaggeration to say that most ordinary people probably cannot understand it.
If you just want to have some fun, try Sora and Veo. However, these websites are not only expensive, but the results are also like lottery tickets. You have to pay for each attempt, making it difficult to use even for people in China.
Who would have thought that after people had been struggling for so long, China’s ByteDance was quietly preparing a big surprise?
(Image source: Jimen)
Just this week, Seedance 2.0, ByteDance’s video model, suddenly became public. There was no long waiting list for applications, no invitations for hidden internal tests. It was only opened to the public during the Spring Festival, the biggest traffic window of the year.
Having used it, all I can say is that good days are ahead for my friends who want to create their own AI videos.
15 seconds of video generation, 1 hour of queuing
First, let’s talk about how to use it.
Seedance 2.0 is now available on the Jimeng platform. Currently, members (at least 69 yuan) can use the latest model directly and can access it from both the web version on their computer and the mobile app. It is expected to be fully available to all users within the next few days.
If you don’t want to pay, you can also use ByteDance’s Xiaoyunque. Currently, new users who log in are given 3 free opportunities to generate videos using Seedance 2.0 and are awarded 120 points each day.
After you exhaust the free opportunities, it will cost 8 points per second to generate a video with Seedance 2.0. This means you can generate up to 15 seconds of video content every day for free, which is enough for a taste test.
(Image source: Lei Technology)
Now let’s take a look at its features.
As you know, most video models in China could only produce silent videos before. Even ByteDance only added voiceover to their Seedance 1.5 version late last year.
Seedance 2.0’s sound and images are now fully integrated.
This new model can generate matching sound effects and background music while creating videos, and supports lip-syncing and emotion matching. When a character speaks, it ensures that the mouth shape is correct and the facial expressions and tone match.
To test its capabilities, I typed in a simple prompt. From a first-person perspective, you sit by the window of an old-fashioned green train, watching the fields outside the window flash and the glass on the table vibrate slightly.
Perhaps because there were so many people who wanted to experience it, we had to wait in line for over an hour before the video was actually generated.
To be honest, I wasn’t surprised by the level of detail in the photos. It was the sound that really gave me goosebumps. In the video, you could feel not only the gentle background music but also the unique bass rhythm of the train running on the tracks. Even as the camera panned over the glass on the table, the ripples in the water caused by the vibrations were clearly visible.
Looking out the window at the fields and the sunset, it’s hard to believe this never actually existed.
This “original sound” experience is completely different from voiceovers added later. This shows that AI is not just creating images. It understands what’s going on in the image and knows what sounds should be made in that environment.
This is very interesting.
But that’s not enough. Having good sound is not enough. Video must also be stable.
Previously, the most feared thing when creating videos using AI was the “plastic surgery” of the characters. One moment, the protagonist is a tough-looking Westerner, and the next, he’s a young Japanese heartthrob. This problem was especially noticeable in scenes with large-scale movement.
To test the consistency of Seedance 2.0, we intentionally increased the difficulty and generated a video of “Alley Fight Between Two Martial Artists in a Puddle on a Rainy Night.”
The theme of the video is goat vs goat.
The results were quite surprising. In fight scenes longer than 10 seconds, the facial features of the two characters were consistent. Even when he did flying kicks and changed his position, the texture of his clothes and the contours of his face remained intact.
There was still a slight smearing effect on some extremely blurry motion frames, but this is a qualitative improvement compared to the previous generation model where the character’s face changed every three seconds.
In terms of basic qualities, Seedance 2.0 is already a very easy to use tool.
One person can go from text to finished video, but there are still audio-text mismatches and image glitches.
Once the basic test is successful, you can increase the difficulty level.
After all, most of our friends who want to do self-media want AI to not only create realistic photos but also understand our creative ideas.
For this reason, Seedance 2.0 introduced a concept called . Selfie storyboard and selfie camera movement.
Simply put, it can automatically plan your storyboard and camera movements based on your description. Just tell us your requirements and we’ll decide how to shoot.
Xiao Lei tried to type in a very simple instruction. A person wearing athletic shoes is running hard on the soft sand as the sun sets.
The difficulty in explaining this lies not only in the storyboard, but also in understanding the physical world.
Sand is a fluid, so when you step on it, you sink, and when you raise your foot, you carry grains of sand with you. These are details that were difficult to reproduce with previous video generation technologies.
In the video produced, you could clearly see the feet digging into the sand. Every time a person pushes the ground, the sand grains fly backwards, and the parabola of flying sand looks very natural. There was no anti-gravity phenomenon in which sand floated in the air. There was also a clear tremor in the way his calf muscles swayed to the rhythm of his running.
To be honest, when I saw this result, the following thought crossed my mind: This effect can be used directly in short videos.
Based on this effect, Can I create a 60 second Brain Rot short video using the workflow directly?
So I first discovered Doubao, another AI assistant from ByteDance. I asked them to generate a rough 9-grid video storyboard as per their requirements and create a very standard Brain Rot short video script on the theme ‘Choose the red door or the blue door’.
(Image source: Lei Technology)
I have to complain that Doubao still doesn’t really understand the storyboard pictures and it took quite a while.
Next, I entered the storyboard and script into Seedance 2.0.
Seedance 2.0 currently only supports a maximum video length of 15 seconds, but through multimodal input, You can use the end of the previous video as material for the next video requirement Complete the connections between multiple shots, keep the characters consistent, then perform manual editing and merging.
This whole process took half a day to complete.
Although the level of Seedance 2.0’s Chinese generation is far above the level of foreign competitors, in the actual generated content, there are cases where subtitles do not match the audio, and text glitches in images are objectively present and almost unavoidable.
Due to the current 15 second limit, if you provide more text content, the synthesized voice will read the entire text at a very unnatural high speed.
Also, the video generated this time was relatively long. You can clearly see that Seedance 2.0 always handles door opening actions in a strange way. Even after using up all my free credits, I couldn’t get any better results and had to give up.
Regarding the “lottery” issue…at least with current video generation applications, this is unavoidable.
conclusion
In my opinion, the arrival of Seedance 2.0 is like a shotgun for domestic creators.
In terms of pure technical metrics and content output, there’s no denying that Sora may still be the industry benchmark in terms of consistency in long shots and artistic flair in photography.
(Image provided by Sora)
But in the tech industry, there’s a very simple truth. Good technology must first be usable.
For now, Seedance 2.0 has few usage thresholds. Anyone can easily register and use it, and the price is very cost-effective compared to similar competitors.
Tim, a well-known self-media blogger from Movie and TV Hurricane, also praised the results produced by today’s Seedance 2.0 model. The company praised the high definition of the video it generates, the movement of the camera, the continuity of storyboards, and the degree of matching between audio and video, calling it “AI that will change the video industry.”
In a sense, the opinions of video industry experts are far more important than the ratings of our own media or the scores of major device rankings.
Over the next six months, you’ll see a number of short dramas, mystery commentaries, and even product promotional videos produced by Seedance 2.0 on your Douyin and video accounts. Content that does not require complex acting skills and focuses on visual spectacle and plot twists will be the first area to be completely transformed by AI.
Can you believe I have no experience in art, animation, or video production?
