New AI algorithms enable faster, more reliable learning

Machine Learning


A recent study by Northwestern Engineering researchers reveals a new artificial intelligence (AI) algorithm tailored for smart robotics. These findings show that Natural machine intelligence. By facilitating the rapid and reliable learning of complex skills, this new method aims to increase the utility and safety of robots in a variety of applications such as self-driving cars, delivery drones, domestic assistants, and automation. is.

New AI algorithms enable faster, more reliable learning
Todd Murphy.Image courtesy of Northwestern University

The algorithm, known as MaxDiff RL, is successful because it motivates the robot to explore its surroundings as haphazardly as possible to gain experience. This “designed randomness” enhances the quality of information the robot collects about its environment.

The simulated robots demonstrated faster and more effective learning using high-quality data, resulting in improved overall performance and reliability.

Northwestern's new algorithm produced robots that consistently simulated better than state-of-the-art models when tested against other AI platforms.

The effectiveness of the new algorithm is remarkable, as the robot can quickly pick up new tasks and successfully execute them on the first try. This represents a significant departure from current AI models that rely on slow trial-and-error learning methods.

Other AI frameworks may be slightly less reliable. Sometimes we complete a task perfectly, and other times we completely fail. With our framework, you can expect your robot to do exactly what you ask it to do every time you turn it on, as long as it can solve the task. This makes it easier to interpret the robot's successes and failures. This is critical in a world that increasingly relies on AI..

Thomas Berrueta, Principal Investigator, Northwestern University

Berueta holds a PhD in mechanical engineering. She is a candidate in the McCormick School of Engineering and a Chancellor's Fellow at Northwestern University. The paper's lead author is Todd Murphy, a robotics expert, professor of mechanical engineering at McCormick College and Belueta's advisor. Dr. Alison Pinoski is a candidate in Murphy's lab and co-author of the paper with Berueta and Murphy.

disembodied amputation

Scientists, developers, and researchers use vast amounts of human-curated and filtered big data to train machine learning algorithms. The AI ​​acquires knowledge from this training set through trial and error and ultimately achieves the optimal result. This method doesn't work for bodied AI systems like robots, but it works well for disembodied systems like ChatGPT and Google Gemini (formerly known as Bard). Instead, robots collect data on their own without the help of human curators.

Traditional algorithms are incompatible with robotics in two different ways. First, intangible systems can take advantage of a world where the laws of physics do not apply. Second, individual failures have no consequences. All that matters for a computer science application is that it succeeds most of the time.In robotics, a single failure can have devastating effects..

Todd Murphy, McCormick School of Engineering Professor of Mechanical Engineering, Robotics Specialist, and senior study author

Murphy is an advisor to Belueta.

To fill this gap, Berrueta, Murphey, and Pinosky set out to create a new algorithm that ensures robots collect high-quality data while moving. Essentially, MaxDiff RL instructs the robot to move more randomly to collect comprehensive and diverse data about its surroundings. Robots learn through random, self-selected experiences and acquire the capabilities needed to perform real-world tasks.

Understand it correctly even the first time

The researchers tested the new algorithm against the most advanced models available at the time. Researchers used computer simulations to train a simulated robot to perform several everyday tasks. In general, robots using MaxDiff RL learn new skills faster than robots using other models. They also completed tasks much more consistently and reliably than others.

More notably, robots utilizing the MaxDiff RL method frequently achieved accurate task execution in a single trial, even when starting with no prior knowledge.

Our robots were faster, more agile, and were able to effectively generalize what they learned and apply it to new situations. This is a huge benefit for real-world applications where robots cannot spend endless hours on trial and error..

Dr. Berueta Candidate and Presidential Fellow, Department of Mechanical Engineering, McCormick School of Engineering

MaxDiff RL is a general algorithm suitable for a variety of applications. The researchers hope that by addressing the fundamental problems holding back the field, reliable decision-making in smart robotics will become possible.

Alison Pinoski added:This doesn't have to only be used for robotic vehicles that move around. It could also be used in stationary robots, such as a robotic arm in the kitchen that learns how to load a dishwasher. As tasks and physical environments become more complex, the role of embodiment becomes even more important to consider during the learning process.This is an important step towards real systems that perform more complex and interesting tasks

single swimmer fixed

A single robot simulation to test new AI algorithms. Video credit: Northwestern University.

Reference magazines:

AT, Belueta, other. (2024) Maximal Diffusion Reinforcement Learning. nature machine intelligence. doi.org/10.1038/s42256-024-00829-3.

Source: https://www.northwestern.edu/



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *