
Important things to know:
- Smartphone Video → 3D Twins: Cornell's drawers turn regular clips into interactive photorealistic rooms with movable doors, drawers and objects.
- Low Friction Capture: There are no special sensors or manual mappings. Casual videos provide real-time geometry, textures, and articulations (hinges, slides).
- Beyond the demo: The team trained robots in a digital kitchen, transferred skills to the real world, pointing to faster, safer automation workflows.
- What's next: Soft/reflective material and support for larger indoor and outdoor scenes is planned. Performance, reliability and edge case handling remain a proactive issue.
Despite years of bold promises and heavy investments, Augmented Reality (AR) remains a technology that struggles to bridge the gap between imagination and implementation. From clunky hardware to complex software challenges, the dream of seamlessly blending the digital and physical worlds is harder to come true than many expect.
Now, Cornell University research teams, together with collaborators from the University of Illinois and the University of Washington, may have taken a major step in new AI systems that create interactive 3D environments besides regular smartphone video. The breakthrough, officially announced at CVPR 2025 and reported by the Cornell Chronicle, has attracted attention in both the gaming and robotics communities.
Why has AR been shortage up until now and why is this breakthrough different? How can this change the way robots interact with the real world?
The expanded world challenge
Augmented reality (AR) has been hailed as a future for at least a decade, but it still feels more like a technological demonstration than a serious tool. Despite billions of investments and some very flashy marketing, AR has struggled to gain serious traction between consumers and developers. But it's not just the issue where hype doesn't live up to expectations. The underlying challenges are technical and practical, and deeper than most people notice.
Starting with the obvious, technology may not be ready yet. Of course, there are smartphones that place digital furniture in the living room and turn your face into something humorous, but that's far from the seamless, immersive augmentation of reality that science fiction has promised. Simply put, the current generation of hardware and software is not simply left to the job.
It is also possible that AR is ahead of its time. Some technologies need to keep up before they make sense, just as it took mobile computing and electric vehicles to become mainstream. At AR, we still think about where actual value goes beyond novelty. Training, navigation and industrial repairs seem like useful applications, but social media goggles and digital billboards floating on your eyeliner are not exactly killer apps.
Then there is a usability issue as AR is notoriously troublesome to use in public places. The headset is bulky, making calls to “see” things through the screen, boring and battery life tank, a matter of context. We still have a long way to go as AR feels natural, intuitive and socially acceptable.
But zooming in further, you bump into the real brick wall, engineering. Building an expanded world that is accurate, interactive and fast enough for real-time use is technically brutal. One of the most difficult problems is mapping the real world to a virtual world.
Lidar (light detection and range) is often used for this purpose. While this is a great technology, Lidar is cheap enough, power efficient or compact enough to narrow it down to all consumer-grade devices. Even when you do, you are still limited by resolution, speed and environmental conditions.
Real-world spaces are complex and dynamic, and the environment itself doesn't work well, as they are full of edge cases. Reflective surfaces, moving people, low-light conditions, messy rooms all interfere with spatial tracking and understanding of the scene.
Finally, there is an elephant in the room: processing power. Creating a compelling AR experience in real time means scanning your environment simultaneously, building 3D maps, recognizing objects, reading motion, rendering graphics, maintaining frame rates, all passively cooled, maintaining frame rates on a mobile chip. It's a lot to ask, and without the acceleration of dedicated hardware, performance quickly becomes a bottleneck and with it becomes a user experience.
Researchers create AI that builds augmented reality from vision
Developed by researchers at Cornell University, the new AI system transforms everyday smartphone video into an interactive, photorealistic 3D world, opening up exciting possibilities for gaming, robotics and digital design. Technology called drawers allows you to shoot simple videos of rooms like kitchens and offices, instantly build detailed digital twins that allow you to effectively open and move objects like drawers and cabinet doors, creating a truly immersive experience.
Unlike previous methods that require specialized equipment or manual mapping, the drawer only requires short casual videos captured on a regular smartphone. One student and lead developer at PhD Cornell, Hong Kong Xia, explained that the system analyzes visual data from the video to reconstruct the shapes, textures and dimensions of the room. This allows AI to create realistic-looking 3D models and naturally respond to user interactions. This is a greater advance over existing technologies that generate static or non-operational models.
Project leader Assistant Professor Wei-chiu Ma stressed that while many current approaches can integrate scenes from different camera angles, they often lack true dialogue. Drawers break new ground by understanding fine geometric details and how parts of the environment can move and how they should move. The system's articulation module detects elements such as horizontally sliding drawers and outward swing doors, while the perceptual module identifies the hinge's location and articulation type. Use the Generated AI to imagine hidden areas like inside the cabinet and enhance realism.
See how the drawer works
It helps you see how you can convert your regular kitchen video into a responsive, interactive digital twin to fully understand the functionality of your drawers. The following demonstrations provided by the research team highlight the system's ability to reconstruct an environment in real time using mobile objects.
Integrating these features into a seamless, functional framework took several months of development. Xia demonstrated the diversity of drawers by recreating a variety of rooms, including her own office and bathroom, when working. In one demonstration, the digital kitchen was transformed into a simple gaming environment using the Unreal engine, where users interact dynamically by throwing virtual balls and knocking over objects, demonstrating the system's natural response to movement.
Beyond the game, drawers offer serious possibilities for robotics. Robots can be trained within these digital replicas through a process called Real-to-Sim-to-Real Transfer, where tasks are practically learned before they are applied in the real world. The MA team managed to train the robot arms in the digital kitchen to clean up the objects in the drawers, and the robots performed the same task accurately in physical space. This method significantly reduces the costs and risks of physical robot training, making automation safer and more accessible.
Looking ahead, researchers plan to expand the drawer functionality beyond hard objects to include soft or deformable materials such as fabrics, as well as reflective and dynamic surfaces such as mirrors and windows. It is also working to expand technology to model the entire building and outdoor environment, envisaging urban planning, smart farming and disaster response applications where accurate digital twin virtual testing can improve decision-making.
How can these technologies change the future of robotics?
The ability to create a completely interactive, photorial environment other than smartphone video is more than just novel. It could represent a major change in how robots perceive and act in the real world. Systems like drawers not only capture geometry but also extract meaning, allowing machines to build practical mental models of environments with object relationships, constraints and potential interactions.
Such concepts represent a major shift in the field of robotics and environmental interactions in general. Traditionally, robots have “referenced” via sensors such as cameras and Lidar, but they are only minimal in practice, and are often limited to bounding boxes and depth maps. The AI-driven scene reconstruction allows robots to recognize that drawers are not merely obstacles, but that open and hold objects and follow certain rules of movement.
This concept unlocks serious new possibilities in robotics planning. This is because the robots that generate digital twins around them can simulate the intended action before committing to them. For example, a robot can effectively test whether there is clearance to open the door, or whether the gripper fits in a narrow cabinet. This kind of real-time self-generated foresight brings robotics closer to something similar to human foresight. Not only will you react to the world, but you will predict it.
The use of this technology goes far beyond industrial weapons and smart homes, with one typical example being autonomous vehicles. In this case, self-driving cars could benefit from a richer environmental model. Instead of simply identifying lanes or cars, the vehicle can interpret subtle features such as the intention to reach the car door or the path of partially obscure cyclist movement, resulting in better decisions.
Of course, this level of refinement is not free. Technology is still evolving, with many challenges remaining, particularly with regard to processing power, reliability and edge case operation. However, the direction is clear. Tools like drawers focus on a future where robots don't just look at the world. They understand that. And that understanding can be key to making intelligent machines practically useful.
