2025 Reasoning and Planning Apple Workshop

Machine Learning


Reasoning and planning are the foundation of intelligent AI systems, allowing them to plan, interact, adapt, and ultimately operate independently. At Apple, understanding and developing the inference capabilities of AI systems has been an active area of ​​research for many years, resulting in numerous publications that explore new techniques to advance the inference frontier as well as further the field’s understanding of the capabilities (and limitations) of current approaches.

Last year, Apple hosted the Reasoning and Planning Workshop, a two-day event that brought together Apple researchers and members of the broader research community to highlight cutting-edge technological advances in this field. This workshop focused on three main areas: inference and planning, application to agents, and model development.

Presentations and discussions of these topics examined how reasoning and planning models process and complete complex tasks based on simple instructions. Workshop participants discussed questions such as:

  • How to ground reasoning processes in different modalities and instantiations.
  • What is the interaction between search/test time computing and underlying model functionality?
  • How can multiple inference systems work together?

The group also considered architectures that leverage memory and adaptation, and how to plan and reason in a reliable, secure, and efficient manner. In addition, workshop participants discussed model development for inference systems using environments and simulators, along with data generation techniques specifically designed for scalable training and reliable benchmarking. Advances in these research areas enable the construction of adaptive, efficient, and intelligent systems that can tackle complex and dynamic challenges.

In this post, we will share recordings of selected talks and summaries of publications discussed at the workshop.

2025 Reasoning and Planning Apple Workshop Video

Published works presented at the workshop

AbstRaL: Enhancing LLM Reasoning by Enhancing Abstract Thinking Written by Silin Gao (research done while at Apple), Antoine Bosselut (EPFL), Samy Bengio, Emmanuel Abbe

Adaptable logic control for large-scale language models by Honhua Zhang (UCLA), Po-Nien Kung (UCLA), Masahiro Yoshida, Guy Van den Broeck (UCLA), Nanyun Peng (UCLA)

“Adapt On-the-Go: Behavior Modulation for Single Life Robot Deployment” by Annie S. Chen (Stanford University), Govind Chada (Stanford University), Laura Smith (University of California, Berkeley), Archit Sharma (Stanford University), Zipeng Fu (Stanford University), Sergey Levine (University of California, Berkeley), Chelsea Finn (Stanford University)

Common sense reasoning about legged robot adaptation using visual language models by Annie S. Chen (Stanford University), Alec M. Lessing (Stanford University), Andy Tang (Stanford University), Govind Chada (Stanford University), Laura Smith (UC Berkeley), Sergey Levine (UC Berkeley), Chelsea Finn (Stanford University)

Embodied Agent Interfaces: Benchmarking LLM for Embodied Decision Making by Manling Li (Stanford University, Northwestern University), Shiyu Zhao (Stanford University), Qineng Wang (Stanford University, Northwestern University) Kangrui Wang (Stanford University, Northwestern University), Yu Zhou (Stanford University), Sanjana Srivastava (Stanford University), Cem Gokmen (Stanford University), Tony Lee (Stanford University), Li Erran Li (Amazon), Ruohan Zhang (Stanford University), Weiyu Liu (Stanford University), Percy Liang (Stanford University), Li Fei-Fei (Stanford University), Jiayuan Mao (MIT), Jiajun Wu (Stanford University)

Ferret-UI: Understanding Grounded Mobile UI with Multimodal LLM by Keen You, Haotian Zhang, Eldon Schoop, Floris Weers, Amanda Swearngin, Jeffrey Nichols, yingfei Yang, Zhe Gan

Ferret-UI 2: Mastering the understanding of universal user interfaces across platforms by Zhangheng Li, Keen You, Haotian Zhang, Di Feng, Harsh Agrawal, Xiujun Li, Mohana Prasad Sathya Moorthy, Jeff Nichols, yingfei Yang, Zhe Gan

From Multimodal LLM to Generalist Embodied Agents: Methods and Lessons by Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev

Grounding multimodal large-scale language models in action by Andrew Szot, Bogdan Mazoure, Harsh Agrawal, Devon Hjelm, Zsolt Kira, Alexander Toshev

GSM-Symbolic: Understanding the limits of mathematical reasoning in language models by Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Onsel Tuzel, Samy Bengio, Mehrdad Farajtabar

How far can Transformers deduce? “The Barrier of Globality and the Guiding Scratchpad” by Emmanuel Abbe, Sammy Bengio, Ario Lotfi, Colin Sandon, and Omid Salemi

“Illusions of Thinking: Understanding the Strengths and Limitations of Inference Models through the Lens of Problem Complexity” by Parshin Shojaee (work done during internship at Apple), Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, Mehrdad Farajtabar

Large-Scale Language Models as Generalizable Policies for Embodied Tasks Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Walter Talbott, Rin Metcalf Susa, Natalie Macraz, Devon Hjelm, Alexander Toshev

Learning Adaptive Parallel Inference with Language Models by Jiayi Pan (UC Berkeley), Xiuyu Li (UC Berkeley), Long Lian (UC Berkeley), Charlie Snell (UC Berkeley), Yifei Zhou (UC Berkeley), Adam Yala (UC Berkeley, UCSF), Trevor Darrell (UC Berkeley), Kurt Keutzer (UC Berkeley), Alane Suhr (UC Berkeley)

Learn to reason without external rewards by Xuandong Zhao (University of California, Berkeley), Zhewei Kang (University of California, Berkeley), Aosong Feng (Yale University), Sergey Levine (University of California, Berkeley), Dawn Song (University of California, Berkeley)

LLM Metacognitive Skills: Exploring Mathematical Problem Solving by Aniket Didolkar (Mila, University of Montreal), Anirudh Goyal (Mila, University of Montreal), Nan Rosemary Ke (Google DeepMind), Siyuan Guo (The University of Cambridge), Michal Valko (Google DeepMind), Timothy Lillicrap (Google DeepMind), Danilo Rezende (Google DeepMind), Yoshua Bengio (Mila, University of Montreal), Michael Mozer (Google DeepMind), Sanjeev Arora (Princeton University)

Mind2Web 2: Evaluating agent search with agents as judges by Boyu Gou (The Ohio State University), Zanming Huang (The Ohio State University), Yuting Ning (The Ohio State University), Yu Gu (The Ohio State University), Michael Lin (The Ohio State University), Weijian Qi (The Ohio State University), Andrei Kopanev (The Ohio State University), Botao Yu (Ohio State University) Bernal Jimenez Gutierrez (Ohio State University), Yiheng Xu (Ohio State University), Zhang Hee Song (Ohio State University), Jaman Wu (Ohio State University), Xijie Chen (Ohio State University), Hanane Noor Moussa (Ohio State University), Tianxu Zhang (Ohio State University), Jiang Xue (Ohio State University), Yifei Li (Ohio State University), Tianci Xue (The Ohio State University), Zeyi Liao (The Ohio State University), Kai Zhang (The Ohio State University), Boyuan Zheng (The Ohio State University), Zhaowei Cai (Amazon AGI), Viktor Rozgic (Amazon AGI), Morteza Ziyadi (Amazon AGI), Huan Sun (The Ohio State University), Yu Su (The Ohio State University)

MMAU: Comprehensive benchmarking of agent capabilities across diverse domains by Guoli ying, Haoping Bai, Shuang Ma, Feng Nan, Yanchao Sun, Zhaoyang Xu, Shen Ma, Jiarui Lu, Xiang Kong, Aonan Zhang, Dian Ang Yap, Yizhe Zhang, Karsten Ahnert, Vik Kamith, Mathias Berglund, Dominic Walsh, Tobias Ginderle, Jurgen Wiest, Lai Zhenfeng, George Horrell, Wang Xiaoming, Jiulong Shan, Meng Cao, Pang Ruiming, Wang Rui

Omega: Can LLM reason outside the box in mathematics? Evaluating exploratory, constructive, and transformative generalization: Yiyou Sun (University of California, Berkeley), Shawn Hu (dmodel.ai), Georgia Zhou (University of California, Berkeley), Ken Zheng (University of California, Berkeley), Hannaneh Hajishirzi (Ai2, University of Washington), Nouha Dziri (Ai2), Dawn Song (University of California, Berkeley)

Understanding the modeling capabilities of large-scale language models for sequential decision making by Martin Klisarov, Devon Hjelm, Alexander Toshev, and Bogdan Mazoure

OSWorld: Benchmarking multimodal agents for open-ended tasks in real computing environments by Tianbao Xie (University of Hong Kong), Danyang Zhang (University of Hong Kong), Jixuan Chen (University of Hong Kong), Xiaochuan Li (University of Hong Kong), Siheng Zhao (University of Hong Kong), Ruisheng Cao (University of Hong Kong), Toh Jing Hua (University of Hong Kong), Zhoujun Cheng (University of Hong Kong), Dongchan Shin (University of Hong Kong), Fangyu Lei (University of Hong Kong), Yitao Liu (University of Hong Kong), Yiheng Xu (University of Hong Kong), Shuyan Zhou (Carnegie Mellon University), Silvio Savarese (Salesforce Research), Caiming Xiong (Salesforce Research), Victor Zhong (University of the University of Waterloo), Taoyu (University of Hong Kong)

Policy learning from tutorial books through understanding, rehearsal, and reflection by Xiong-Hui Chen (Nanjing University), Ziyan Wang (King’s College London), Yali Du (King’s College London), Shengyi Jiang (University of Hong Kong), Meng Fang (University of Liverpool), Yang Yu (Nanjing University), and Jun Wang (University College London).

RAGEN: Understanding Self-Evolution of LLM Agents with Multiturn Reinforcement Learning by Zihan Wang (Northwestern University), Kangrui Wang (Northwestern University), Qineng Wang (Northwestern University), Pingyue Zhang (Northwestern University), Linjie Li (University of Washington), Zhengyuan Yang (Microsoft), Xing Jin (University of British Columbia), Kefan Yu (Northwestern University), Minh Nhat Nguyen (Singapore Management University), Licheng Liu (Northwestern University), Eli Gottlieb (Northwestern University), Yiping Lu (Northwestern University), Kyunghyun Cho (New York University), Jiajun Wu (Stanford University), Li Fei-Fei (Stanford University), Lijuan Wang (Microsoft), Yejin Choi (Stanford), Manlin Lee (Northwestern University)

Reinforcement Learning for Long-Horizon Interactive LLM Agents by Kevin Chen, Marco Cusumano-Towner, Brody Huval, Aleksei Petrenko, Jackson Hamburger, Vladlen Koltun, Philipp Krähenbühl

Tracking from the future: A probabilistic reasoning approach to controllable language production Gwen Yidou Weng (UCLA), Benjie Wang (UCLA), Guy Van den Broeck (UCLA)

Training Software Engineering Agents and Verifiers with SWE-Gym by Jiayi Pan (UC Berkeley), Xingyao Wang (UICU), Graham Neubig (CMU), Navdeep Jaitly (Apple), Heng Ji (UIUC), Alane Suhr (UC Berkeley), Yizhe Zhang (Apple)

Upsample or Upweight? Balanced Training on Significantly Imbalanced Datasets by Tianjian Li (Johns Hopkins University), Haoran Xu (Johns Hopkins University), Weiting Tan (Johns Hopkins University), Kenton Murray (Johns Hopkins University), Daniel Khashabi (Johns Hopkins University)

Cheers to robots: Improving language on the fly with language correction by Lucy Xiaoyang Shi (Stanford University), Zheyuan Hu (University of California, Berkeley), Tony Z. Zhao (Stanford University), Archit Sharma (Stanford University), Karl Pertsch (Stanford University, University of California, Berkeley), Jianlan Luo (University of California, Berkeley), Sergey Levine (University of California, Berkeley), Chelsea Finn (Stanford University)

Many people contributed to this workshop, including Akshay Aggarwal, Harsh Agrawal, Felix Bai, Samy Bengio, Meng Cao, Mehrdad Farajtabar, Philipp Krähenbühl, Iman Mirzadeh, Yanchao Sun, Alexander Toshev, Simon Wang, Guoli ying, and Keen You.



Source link