It seems like every day there's a new AI video announcement, and the latest comes from Hedra, a startup that's taking a character-first approach to turning ideas into action.
This week alone saw the announcement of new features for the Luma Labs Dream Machine and the new Sora-like Gen-3 from Runway.
Character-1 is a research preview of an upcoming foundational video model that will give users fine-grained control over how to animate their virtual characters using AI.
Preview takes your audio and images and creates a lip-sync video of the character in the image speaking. Unlike other lip-sync tools, this one adds advanced expressions and movements never seen before.
During this research preview, Hedra is free and allows users to create videos of any length, which the company is using to test issues in both its models and moderation tools before rolling out more advanced features.
How does Hedra Character-1 work?
Character-1 is a new foundational AI model designed to use AI to create fully controllable, realistic characters that the company says can speak expressively, sing, and even rap at infinite length.
For now, it's pretty easy to use. Once you've signed up, you can generate audio from text or give your own audio to create a character. This can be generated from a photo, an AI image, or text, and generate the image within Hedra. After that, you just click generate video and wait.
There are similarities in functionality to lip-sync tools from several open source projects, research previews, and even platforms like Runway and Synclabs, but what sets Hedra apart for me is its future-proofing and expressive power for video.
Speaking about its future plans, the company said: “This is the first step in Hedorah's mission to build a multi-modal creation studio that is accessible to everyone, giving creators full control over emotional interactions, movement, and (yes) the entire world.”
How well does Hedra Character-1 work?
Introducing a research preview of our foundational model, Character-1. Available today at https://t.co/G45zFlUfcN (desktop and mobile). * Unlimited duration (30 seconds for open preview)* 90 seconds generated every 60 seconds (while H100 supply lasts)* Expressive speaking, singing, rapping… pic.twitter.com/cYuHpSnqMuJune 18, 2024
Because this is the first phase of a new model, there have been some initial issues, especially with the overly strict moderation AI, but I haven’t had any issues with the videos it generated.
There's currently a 30 second limit, so if you have a longer audio clip like me, you'll need to split it into two sections. It works best to use an image generated by Hedra, but you can also upload your own image – just make sure it faces forward, like a human.
For now, only square format video is offered (not widescreen or portrait) and the resolution is relatively low, but this is a research preview to showcase features rather than produce production-ready content, and it really is a glimpse into what's to come.
To test it, I created a short story about an alien invasion, which generated four characters: three aliens from the galactic fleet and one human general. Compared to the human performances, it has the awkwardness of a student soap opera, but for AI-based lip syncing, it's a vast improvement over what we've seen before.
