EMBridge: Enhanced gesture generalization from EMG signals by cross-modal representation learning.

Machine Learning


Hand gesture classification using high-quality structured data such as videos, images, and hand skeletons is a well-studied problem in computer vision. Alternatively, low-power, cost-effective biosignals such as surface electromyography (sEMG) can be leveraged to enable continuous gesture prediction on wearable devices. In this work, we aim to improve the quality of EMG representations and ultimately enable generalization of zero-shot gestures by adjusting EMG representations with embeddings obtained from structured, high-quality modalities that provide richer semantic guidance. Specifically, we propose EMBridge, a cross-modal representation learning framework that bridges the modality gap between EMG and posture. EMBridge learns high-quality EMG representations by introducing a Querying Transformer (Q-Former), a masked pose reconstruction loss, and a community-aware soft-contrast learning objective that adjusts the relative geometry of the embedding space. We evaluated EMBridge on both in-distribution and invisible gesture classification tasks and demonstrated consistent performance improvements across all baselines. To the best of our knowledge, EMBridge is the first cross-modal representation learning framework to achieve zero-shot gesture classification from wearable EMG signals, demonstrating potential towards real-world gesture recognition on wearable devices.



Source link