Oxford researchers propose Farm3D: an AI framework that can learn articulated 3D animals by extracting 2D diffusions for real-time applications like video games

AI Video & Visuals


https://arxiv.org/abs/2304.10535

The tremendous growth of generative AI has brought compelling advances in image production, with techniques like DALL-E, Imagen, and Stable Diffusion creating great images from text cues. This work may extend beyond 2D data. As recently demonstrated with DreamFusion, you can use a text-to-image generator to create high-quality 3D models. The generator has no 3D training, but has enough data to reconstruct the 3D shape. In this article, I’ll show you how to make more use of the text-to-image generator to get joint models for several 3D item types.

So, instead of trying to create a single 3D asset (DreamFusion), I want to create a statistical model of a whole class of articulated 3D objects (cows, sheep, horses, etc.) that can be used to create animatable 3D. I’m here. Assets that can be used to create content from a single image, whether AR/VR, games, and digitally created. They tackle this problem by training a network that can predict a concatenated 3D model of an object from his single photo of the object. To deploy such reconstruction networks, previous efforts have relied on real data. However, we suggest using synthetic data generated using a 2D diffusion model such as Stable Diffusion.

Researchers at the University of Oxford’s Visual Geometry Group use 3D generators such as DreamFusion, RealFusion and Make-a-video-3D, along with test-time optimization to create a single static or dynamic 3D asset I’m proposing Farm3D. , starting with text or images, takes hours. This has several advantages. To begin with, 2D image generators tend to produce accurate, pristine examples of object categories, implicitly curating training data and streamlining learning. An even clearer understanding is provided by the 2D generator implicitly providing a virtual view of each object instance given by the distillation. Third, it removes the need to collect (and possibly censor) the actual data, making the approach more adaptable.

πŸš€ Check out 100 AI Tools in the AI ​​Tools Club

During testing, their network performed reconstructions in seconds from a single image in a feed-forward manner, and instead of fixed 3D or 4D artifacts, they created articulated images that could be manipulated (animated, relighted, etc.). Generate a 3D model. Their method is suitable for synthesis and analysis. This is because the reconstruction network generalizes to real images and is trained on virtual inputs only. Applications can be created to study and store animal behavior. Farm3D is based on two key innovations. To learn concatenated 3D models, we first show how to use rapid engineering to induce stable diffusion to generate a large training set of globally clean images of object categories.

They show how MagicPony, a state-of-the-art technique for monocular reconstruction of articulated objects, can be bootstrapped using these photographs. Then, instead of fitting a single radiance field model, we extend the Score Distillation Sampling (SDS) loss to achieve synthetic multiview surveillance and use a photogeometric autoencoder (in their case his MagicPony) It shows that you can train. To create new artificial views of the same object, photogeometric autoencoders divide the object into various aspects that contribute to image formation (object articulation, appearance, camera viewpoint, lighting, etc.).

These synthetic views are fed into the SDS loss to obtain gradient updates and backpropagation to the learnable parameters of the autoencoder. They provide Farm3D with a qualitative assessment based on its 3D creation and restoration capabilities. Farm3D can be reconfigured as well as created, so it can be quantitatively evaluated in analytical tasks such as semantic keypoint transfer. The model does not use real images for training, saving time-consuming data collection and curation, but performs as well or better than various baselines.


check out paper and plandon’t forget to join 20,000+ ML SubReddit, cacophony channeland email newsletterWe share the latest AI research news, cool AI projects, and more. If you have any questions about the article above or missed something, feel free to email me. Asif@marktechpost.com

πŸš€ Check out 100 AI Tools in the AI ​​Tools Club

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing a Bachelor’s Degree in Data Science and Artificial Intelligence from the Indian Institute of Technology (IIT) in Bhilai. He spends most of his time on projects aimed at harnessing the power of machine learning. His research interest is image processing and his passion is building solutions around it. He loves connecting with people and collaborating on interesting projects.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *