LDM3D is the industry’s first generative AI model to provide depth mapping. It has the potential to revolutionize content creation, the metaverse, and digital experiences.
SANTA CLARA, Calif.–(BUSINESS WIRE)– Intel (Nasdaq: INTC):
This press release features multimedia. Read the full release here: https://www.businesswire.com/news/home/20230621842353/en/
Intel Labs, in collaboration with Blockade Labs, introduced the Latent Diffusion Model for 3D (LDM3D). This is the industry’s first diffusion model that offers depth mapping that creates a 3D image of him in a crisp, immersive 360-degree view. LDM3D has the potential to revolutionize content creation, metaverse applications and digital experiences, transforming a wide range of industries from entertainment and gaming to architecture and design. (Credit: Intel Corporation)
what’s new: Intel Labs, in collaboration with Blockade Labs, introduced 3D Latent Diffusion Model (LDM3D), a new diffusion model that uses generative AI to create realistic 3D visual content. LDM3D is the industry’s first model that uses a diffusion process to generate a depth map to create a crisp, immersive 360-degree view of him in 3D. LDM3D has the potential to revolutionize content creation, metaverse applications and digital experiences, transforming a wide range of industries from entertainment and gaming to architecture and design.
“Generative AI technology aims to further extend and enhance human creativity and save time. Only a few models can generate 3D images.Unlike existing latent stable diffusion models, LDM3D allows users to generate images and depth maps from given text prompts with roughly the same number of parameters. This results in more accurate relative depth for each pixel in the image compared to standard post-processing methods for depth estimation, saving developers a lot of time developing scenes. can.”
—Vasudev Lal, Intel Labs AI/ML Research Scientist
Why it matters: A closed ecosystem limits scale. And Intel’s commitment to the true democratization of AI will make the benefits of AI more widely available through an open ecosystem. One area that has seen significant progress in recent years is the area of computer vision, and specifically generative AI. However, many of today’s advanced generative AI models are limited to generating 2D images only. Unlike existing diffusion models that typically generate only 2D RGB images from text prompts, LDM3D allows users to generate both images and depth maps from given text prompts. LDM3D uses almost the same number of parameters as latent stable diffusion to provide more accurate relative depths for each pixel in an image compared to standard post-processing methods for depth estimation .
This research has the potential to revolutionize the way users interact with digital content by allowing them to experience text prompts in ways never thought possible. Using imagery and depth maps generated by LDM3D, users can transform textual descriptions of serene tropical beaches, modern skyscrapers, or sci-fi universes into detailed 360-degree panoramas of her. The ability to capture this depth information instantly increases overall realism and immersion in applications ranging from entertainment and gaming to interior design and real estate, to virtual museums and immersive virtual reality (VR) experiences. enabling innovative applications in diverse industries.
On June 20th, LDM3D won the best poster award at CVPR’s 3DMV workshop.
Usage: LDM3D was trained on a dataset constructed from a subset of 10,000 samples of the LAION-400M database containing over 400 million image-caption pairs. The team used the Dense Prediction Transformer (DPT) deep estimation model (previously developed in Intel Labs) to annotate the training corpus. The DPT-large model provides highly accurate relative depth for each pixel in the image. The LAION-400M dataset is built for research purposes to enable testing of model training at scale for a wide range of researchers and other interested communities.
The LDM3D model is trained on an Intel AI supercomputer powered by Intel® Xeon® processors and Intel® Habana Gaudi® AI accelerators. The resulting model and pipeline combine the generated RGB images with depth maps to produce a 360-degree view for an immersive experience.
To demonstrate the potential of LDM3D, researchers at Intel and Blockade developed DepthFusion, an application that leverages standard 2D RGB photos and depth maps to create immersive and interactive 360-degree viewing experiences. DepthFusion turns text prompts into interactive, immersive digital experiences using TouchDesigner, a node-based visual programming language for real-time interactive multimedia content. The LDM3D model is a single model that creates both the RGB image and its depth map, saving memory footprint and improving latency.
What’s next: The introduction of LDM3D and DepthFusion paves the way for further advances in multi-view generative AI and computer vision. Intel will continue to explore the use of generative AI to build a strong ecosystem of open source AI R&D that augments human capabilities and democratizes access to this technology. Continuing Intel’s strong support for the open ecosystem in AI, LDM3D is open sourced through HuggingFace. This will allow AI researchers and practitioners to further improve this system and fine-tune it for custom applications.
Other context: Intel’s research will be presented at the IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), June 18-22. For more information, seeLDM3D: 3D Latent Diffusion Modelor LDM3D Demo.
About Intel
Intel (Nasdaq: INTC) is an industry leader, creating world-changing technologies that enable global progress and enrich lives. Inspired by Moore’s Law, we continually strive to advance semiconductor design and manufacturing to meet our customers’ biggest challenges. Unlocking the potential of data to transform business and society for the better by embedding intelligence in the cloud, network, edge and computing devices of all kinds. For more information on Intel innovations, visit newsroom.intel.com and intel.com.
View the source version on businesswire.com. https://www.businesswire.com/news/home/20230621842353/en/
Laura Stadler
laura.stadler@intel.com
Source: Intel
Release Date June 21, 2023 • 9:00 AM EDT
