The screenwriter discusses the character’s motivations. The director specifies the camera angle. Editors review takes for consistency and reject clips where gravity appears to be optional or where faces deform mid-scene. After 20 minutes, the full-length music video appears. Singers perform, stories unfold, and visuals sync to every beat. No humans touch the production. The entire film crew was artificial intelligence.
Developed by researchers at Queen Mary University of London and collaborators from four countries, AutoMV is the first open-source system that can generate complete music videos directly from songs. Feed your audio and our expert AI agents work together like a virtual production team to deliver your finished video from start to finish.
This is an ambitious claim in a field littered with false starts. Generative AI has created impressive short clips, but what about long-form storytelling with musical coordination and consistent characters? That’s been stubbornly out of reach. Until now, probably.
The system works by dividing labor between AI agents, each with a different role. First, the music analysis tool analyzes the song, extracts the beat, identifies verse and chorus structure, and transcribes the time-stamped lyrics. Screenwriting agents interpret this data to create scene descriptions and character profiles. The director agent then generates detailed prompts and keyframe images for each shot. A video generator generates a clip, and a verification agent scores it for physical realism and narrative coherence, and requests regeneration if necessary.
Yinghao Ma, the doctoral student who led the study, clearly sees the potential. “Creating a complete music video that follows the entire song was difficult with an AI system,” he says. “We’re especially pleased that this initiative makes music video creation more accessible for independent artists, who now have the ability to share their work on YouTube.”
The economic situation is alarming. Producing a traditional music video takes 40 to 120 hours in a studio and requires a team of 10 or more professionals, including writers, directors, actors, and editors. Costs typically exceed £10,000 per truck. AutoMV brings this down to the price of an API call (probably between £10 and £20). Time investment? Probably about 30 minutes.
For independent musicians working on a limited budget, this is a big change. It could allow solo artists to create visuals to match their sound in their bedroom studios, potentially leveling the playing field long dominated by major label resources.
But is it actually effective? Researchers tested AutoMV against two commercial video generation platforms and evaluated its technical quality, post-production consistency, musical integrity, and artistic merit. Human experts, including music industry experts, record label practitioners, and music video directors, evaluated the works across 12 criteria.
AutoMV significantly outperformed both commercial baselines. This allowed for better character consistency (faces and clothing remained stable across scenes), tighter audio-visual synchronization, and a stronger narrative structure. In terms of musical theme relevance and emotional expression, scores have come closer to those of professionally directed videos, but the quality gap still remains.
Commercial systems have struggled in terms of clarity. One produced a mostly static image with minimal movement. The other created story-driven content but relied on a fixed bank of characters that could not adapt to the input music, resulting in generic casting without regard to the song’s emotional content or cultural context.
AutoMV’s multi-agent architecture addresses these limitations through specialization. A scriptwriting agent not only transcribes the lyrics, but also interprets the semantic meaning and extracts thematic cues and emotional tone. It maintains a “character bank” that stores detailed appearance descriptions such as hair color, age, clothing, and facial features that are preserved between shots. When the director agent generates a prompt, it retrieves these profiles so that the same face is displayed throughout.
The system switches generation approaches depending on shot requirements. Cinematic storytelling uses a single video model. For scenes that require lip-sync precision, route isolated vocal tracks to specialized audio-video conversion models. The final quality control step rejects physically impossible outputs such as floating objects, impossible poses, and glowing eyes before assembly.
Challenges still remain. Dance sequences are offbeat at times. Accurately matching rhythms across diverse musical styles can prove difficult. If your script requires close-ups of handwriting or screen displays, the text will fail to render and become distorted or unreadable. Also, without source separation to separate the vocals, lip sync accuracy will be significantly reduced, especially for songs with complex vocal production.
A research team from Queen Mary, Beijing University of Posts and Telecommunications, Nanjing University, Hong Kong University of Science and Technology, and the University of Manchester has released the system as open source. They invite contributions to the codebase and encourage experimentation with long-form, multimodal AI systems.
Whether AutoMV signals democratization or forced displacement depends on your point of view. Independent musicians gain production capabilities that were previously out of reach. But professional video directors are facing new competition from software that costs a fraction of the daily price. The impact on labor will become apparent over time.
For now, this technology is somewhere in the middle. Enough to be helpful, but not yet enough to replace human creativity at its best. Human-directed videos still score higher on most artistic criteria. The professional cinematography, lighting design, and narrative sophistication are still ahead of the curve. But the gap is closing and the direction forward is clear.
Ma’s team recognizes the ethical complexities. They advocate requiring AI-generated labels on all AutoMV output to maintain the distinction between synthetic and real media. This system does not redistribute copyrighted audio. Researchers access songs via YouTube URLs for evaluation purposes only. Future safeguards could include imperceptible audio watermarks for traceability and community-driven audits to detect abuse.
The broader question is: What happens when such powerful content creation tools become widely accessible? There are millions of music videos on YouTube, many by unsigned artists hoping to be recognized by their audience. If AutoMV delivers on its accessibility promise, that number could explode. Whether your audience values quantity over quality is another matter entirely.
Somewhere out there, an independent musician with talent beyond their budget is probably uploading a song right now. In the near future, complete music videos produced by staff that exist only within the software may also be uploaded. Of course, whether anyone will watch it is a whole other challenge.
Research link: https://arxiv.org/abs/2512.12196
If our reporting has informed or inspired you, please consider making a donation. Every contribution, however big or small, helps us continue to provide accurate, engaging and trustworthy science and medical news. Independent journalism takes time, effort and resources. Your support allows us to continue uncovering the stories that matter most to you.
Join us in making knowledge accessible and impactful. Thank you for your cooperation!
