Generative models of photography are of unprecedented interest due to recent advances in underlying modeling techniques. Today’s most effective models are based on diffusion models, autoregressive transformations and generative adversarial networks. Particularly desirable features of diffusion models (DM) include a resilient and scalable training objective and a tendency to require fewer parameters than their transformer-based counterparts. The lack of large, publicly available general-purpose video datasets and the high computational costs involved in training on video data are the main reasons why video modeling lags behind. At the same time, the field of photography has made great strides.
Although there is a wealth of research on video synthesis, most efforts, including early video DM, produce only low-resolution, often short films. Create enhanced high-definition films by applying video models to real problems. They focus on his two problems related to real-world video generation. (i) text-guided video synthesis for creating creative content; and (ii) video synthesis of high-resolution, real-world driving data with great potential as a simulation engine for autonomous driving. driving. To do this, they rely on the latent diffusion model (LDM). This can significantly reduce the computational load when learning from high-resolution images.
Generate temporally consistent videos using a pre-trained image diffusion model. The model first generates batches of samples that are independent of each other. The samples are temporally aligned to create a coherent film after temporal video fine-tuning.
Researchers from LMU Munich, NVIDIA, Vector Institute, University of Toronto, and University of Waterloo recommend video LDM and extend LDM to high-definition video creation. This is a process that requires a lot of computational power. In contrast to previous work on DM for video creation, their video LDM is initially pre-trained on images only (or uses an existing pre-trained image LDM), resulting in a large number of images. Leverage datasets. After adding the temporal dimension to the latent spatial DM, we transform the LDM image generator into a video generator by modifying the pretrained spatial layers and training only the temporal layers of the encoded image sequence or film (Fig. 1). Establish temporal coherence in pixel space. Tune the LDM’s decoder in a similar way (Figure 2).
It also temporally aligns the pixel space and the potential DM upsampler, which are frequently used for image super-resolution, into a time-consistent video super-resolution model to further improve spatial resolution. Their approach, based on LDM, has the potential to create longer movies that are overall cohesive while using very little memory and processing power. The video upsampler has little training and compute demand as it only needs to work locally to synthesize at very high resolutions. For state-of-the-art visual quality, 5121024 real-world driving scenario films are used to test the technology and synthesize several minutes of video.
In addition, it enhances the powerful text-to-image LDM known as Stable Diffusion, allowing it to be used to create text-to-video at resolutions up to 1280 x 2048. A fairly small training set of movies with captions is available. This is because in such scenarios, a temporal alignment layer needs to be trained. They present the first instance of personalized text-to-video creation by transferring the learned temporal layer to his variably composed text-to-image LDM. They hope their work will pave the way for more effective digital content generation and autonomous driving simulations.
Here are their contributions:
(i) provide a practical method for developing LDM-based video production models with high resolution and long-term consistency; Their key finding is to use pre-trained image DMs to generate videos. By adding a temporal layer, the images can always be aligned consistently (Figures 1 and 2).
(ii) they further refine the super-resolution DM widely used in the timing literature;
(iii) You can create minutes of film and record real-world driving scenarios for state-of-the-art high-definition video compositing performance.
They (i) upgrade the publicly accessible Stable Diffusion text-to-image LDM to a robust and expressive text-to-video LDM (ii), (iii) the learned temporal layer is Indicates that it may be merged with other image model checkpoints (as such). (as DreamBooth), and (iv) do the same for the learned temporal layer.
check out paper and plandon’t forget to join Our 19k+ ML SubReddit, cacophony channeland email newsletterWe share the latest AI research news, cool AI projects, and more. If you have any questions about the article above or missed something, feel free to email me. Asif@marktechpost.com
🚀 Check out 100 AI Tools in the AI ​​Tools Club
Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing a Bachelor’s Degree in Data Science and Artificial Intelligence from the Indian Institute of Technology (IIT), Bhilai. He spends most of his time on projects aimed at harnessing the power of machine learning. His research interest is image processing and his passion is building solutions around it. He loves connecting with people and collaborating on interesting projects.
