Apple researchers share a lot of research through publications and conference engagement to advance AI and ML through basic research, support the broader research community, and help accelerate advancement in the field. This week, the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) will be held in Nashville, Tennessee. Apple is proud to rejoin this important event for the community and become an industry sponsor.
At the main conference and related workshops, Apple researchers present new research across many topics in computer vision, including vision language models, 3D photogrammetry, large-scale multimodal models, and video diffusion models.
CVPR participants will be able to experience a demonstration of Apple's ML research at Booth #1217 during exhibition hours. Apple also sponsors and participates in many Affinity Group Host events that support underrated groups in the ML community. A comprehensive overview of Apple's participation and contributions to CVPR 2025 can be found here. Below are some highlight selections:
FASTVLM: Efficient vision encoding of vision language models
Vision Language Models (VLM) performance improves as the resolution of the input image increases, but popular visual encoders such as VITS become inefficient at high resolution due to the large number of tokens and high encoding latency. In many production use cases, VLMs need to be accurate and efficient to run on devices for AI experiences that meet the low latency demands of real-time applications and provide privacy.
At CVPR 2025, Apple researchers will present efficient vision encoding for the fastVLM: Vision language model. This task shares FastVithd: a new hybrid vision encoder designed to output fewer tokens and significantly reduce encoding times for high resolution images. Using this efficient encoder for high resolution inputs, FastVLM significantly improves the trade-off of precision delays with a simple design. FASTVLM offers accurate, fast and efficient visual query processing, suitable for on-device power supply for real-time applications, with MLX-based inference codes, model checkpoints and iOS/MACOS demo apps available here.
Matrix3d: Large photogrammetry model all-in-one
Photogrammetry allows you to construct 3D scenes from 2D images, but traditional approaches have two limitations. First, a dense collection of 2D images is usually required to achieve robust and accurate 3D reconstruction. Second, pipelines generally involve multiple processes of independent tasks that are uncorrelated or unoptimized with one another, such as functional detection, structure from structure, and multiview stereo.
In a highlight presentation at CVPR, Apple researchers present a new approach to this challenge that overcomes these previous limitations. Paper Matrix3D: Large photogrammetry model A single integrated model that performs several photogrammetry subtasks including all-in-share pose estimation, depth prediction, and synthesis of new views. Matrix3D utilizes a multimodal diffusion transformer (DIT) to integrate transformations across several modalities such as images, camera parameters, and depth maps. Multimodal training in this approach integrates mask learning strategies that allow full modality training even with partially complete data, such as bimodality data for image poses and deep pairs of images, and significantly increases the pool of available training data. Matrix3D presents cutting-edge performance of pose estimation and new view synthesis tasks, providing fine grain control through multi-round interactions, making it an innovative tool for 3D content creation. The code is available here.
Multimodal pre-autoregression training for large vision encoders
Large multimodal models are generally trained by pairing large language decoders with Vision encoders. These vision encoders are usually pre-trained for discriminatory purposes such as contrasting losses, which creates discrepancies between pre-training and the generative autoregressive downstream task. Following the success of the autorecovery approach to train language models, autoreplay image models have been shown to pretrain powerful and scalable vision encoders.
In a highlight presentation at CVPR 2025, Apple ML researchers share multimodal pre-autoregression training for large visual encoders that describe AIMV2, a family of large and powerful visual encoders pre-trained for multimodal autoregression purposes. Multimodal decoder generates both raw patches and text tokens, and these models are excellent for not only multimodal tasks but also visual recognition benchmarks such as localization, grounding, and classification. This work also shows that the AIMV2 model is efficient in training and outperforms current art with significantly fewer samples seen before training. Code and model checkpoints can be found here.
World-Consistent Video Distribution Using Explicit 3D Modeling
Diffusion models have become the dominant paradigm of realistic image and video generation, but these models still struggle to efficiently and explicitly generate 3D consistent content. Traditionally, these methods implicitly learn 3D consistency by generating only RGB frames, leading to training artifacts and inefficiencies.
In the highlight presentation at CVPR, Apple researchers will share video spreads that are globally consistent with explicit 3D modeling detailing new approaches to address these challenges. This technique, World-Consistent Video Dispersion (WVD), trains a diffusion transformer to learn the co-distribution of both RGB (color) and XYZ (space coordinates) frames. As a result, the model can adapt to multiple tasks with flexible im-pinting capabilities. For example, the model can estimate XYZ frames, taking into account ground truth RGB. Alternatively, you can generate a new RGB frame using XYZ projection along the specified camera trajectory. With this flexibility, WVD integrates tasks such as single image to 3D generation, multi-view stereo and camera-controlled video generation.

Demonstration of ML research at the Apple booth
During exhibit hours, CVPR participants will be able to interact with the Apple ML Research live demo of Booth #1217, including the FASTVLM above.
Support for the ML Research Community
Apple is committed to supporting underrated groups in the ML community. We are proud to responsibilities to multiple affinity groups, including multiple affinity groups hosting events in CVPR, including Latinas (LXCV is a subgroup of LXAI) (June 11 workshop), and Women in Computer Vision (WICV) (June 12 workshop).
Learn more about Apple ML Research at CVPR 2025
CVPR brings together a community of researchers promoting cutting edge in computer vision, and Apple is proud to re-share new innovative research at the event and connect with the community that participates in it. This post highlights the selection of works that Apple ML researchers will present at CVPR 2025, and a comprehensive overview and schedule of participation can be found here.
