Using a contact microphone as a tactile sensor for robot manipulation

Machine Learning


Using contact microphones as tactile sensors for robot operation

Two-stage model training. AVID and R3M pre-training leverages large-scale internet video data (blue dashed box). The resulting pre-trained representations are used to initialize vision and audio encoders, and train the entire policy end-to-end by replicating behaviors from a small number of in-domain demonstrations. The policy takes image and spectrogram inputs (left) and outputs a set of actions in delta-end effector space (right). Credit: Mejia et al.

To complete real-world tasks in home environments, offices, and public spaces, robots need to be able to effectively grasp and manipulate a variety of objects. In recent years, developers have created a variety of machine learning-based models designed to enable skilled object manipulation by robots.

Although some of these models have achieved good results, they usually need to be pre-trained on large amounts of data to achieve good performance. The datasets used to train these models consist mainly of visual data, such as annotated images or video footage captured by cameras, although there are also approaches that analyze other sensory inputs, such as tactile information.

Researchers from Carnegie Mellon University and Olin College of Technology recently explored the possibility of using contact microphones instead of traditional tactile sensors, allowing them to use audio data to train machine learning models for robot manipulation. Their paper, posted on a preprint server, states: arXivmay open up new opportunities for large-scale multisensory pre-training of these models.

“Pre-training with large amounts of data is beneficial for robot learning, but current paradigms only perform extensive pre-training of visual representations, while representations for other modalities are trained from scratch,” Jared Mejia, Victoria Dean and their colleagues write in the paper.

“In contrast to the abundance of vision data, relevant internet-scale data that can be used to pre-train other modalities, such as tactile sensing, is unknown. Such pre-training becomes increasingly important in low-data domains, typical in robotic applications. We address this gap by using a contact microphone as an alternative tactile sensor.”







Credit: Mejia et al. (https://sites.google.com/view/hearing-touch)

As part of their recent work, Mejia, Dean, and collaborators pre-trained a self-supervised machine learning approach on audiovisual representations from the Audioset dataset. The Audioset dataset contains over 2 million 10-second video clips of speech and music clips collected from the internet. The pre-trained model relies on audiovisual instance identification (AVID), a technique that can learn to distinguish between different types of audiovisual data.

The researchers evaluated their approach in a series of tests in which they challenged the robot to complete real-world manipulation tasks, relying on up to 60 demonstrations for each task. Their findings were highly promising, with their model outperforming robot manipulation policies that rely solely on visual data, especially when objects and locations differed significantly from those included in the training data.

“Our key finding is that contact microphones inherently capture audio-based information, allowing us to leverage large-scale audio-visual pre-training to derive representations that improve robotic manipulation performance,” Mejia, Dean, and their colleagues write. “To our knowledge, our method is the first approach to leverage large-scale multisensory pre-training for robotic manipulation.”

In the future, the work by Mejia, Dean and colleagues may pave the way for achieving skilled robot manipulation leveraging pre-trained multimodal machine learning models. Their proposed approach may soon be further refined and tested on a wider range of real-world manipulation tasks.

“Future work could explore which characteristics of pre-training datasets are most useful for learning audio-visual representations of manipulation policies,” Mejia, Dean, and colleagues wrote. “Furthermore, a promising direction would be to equip end-effectors with visual-tactile sensors and contact microphones with pre-trained audio representations to determine how to leverage both to provide robotic agents with a deeper understanding of their environment.”

For more information:
Jared Mejia et al. “Auditory-Haptic: Audiovisual Pretraining for Touch-Rich Manipulation” arXiv (2024). Translation: 10.48550/arxiv.2405.08576

Journal Information:
arXiv

© 2024 Science X Network

Quote: Using Contact Microphones as Tactile Sensors for Robot Manipulation (May 30, 2024) Retrieved May 30, 2024 from https://techxplore.com/news/2024-05-contact-microphones-tactile-sensors-robot.html

This document is subject to copyright. It may not be reproduced without written permission, except for fair dealing for the purposes of personal study or research. The content is provided for informational purposes only.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *