
Voice technology has advanced significantly over the past decade, enabling the integration of voice technology into a wide variety of consumer products. Training a good machine learning model for such a job requires a large amount of labeled data, in this case thousands of hours of audio including transcriptions. This information exists only for some languages. For example, of the more than 7,000 languages in use today, only about 100 are supported by current speech recognition algorithms.
Recently, thanks to self-supervised speech representations, the amount of labeled data required to build speech systems has been greatly reduced. Despite progress, current major efforts still cover only about 100 languages.
Facebook’s Massive Multilingual Speech (MMS) project is addressing some of these obstacles by combining wav2vec 2.0 with new datasets containing labeled data in over 1,100 languages and unlabeled data in nearly 4,000 languages. deal with. Based on their findings, a large-scale multilingual speech model outperforms state-of-the-art methods, supporting 10x more languages.
Their initial goal was to collect speech data for hundreds of languages, as the largest speech dataset available only contains up to 100 languages. As a result, they focused on religious documents such as the Bible, which have been translated into many languages and whose translations have been extensively reviewed for text-based language translation research. People recorded themselves reading these translations and made the audio files available online. The study collected New Testament readings in over 1,100 languages, yielding an average of 32 hours of data per language.
Their research showed that the proposed model works equally well for male and female voices, even though this data is from a specific domain and is typically read by male speakers. became clear. Even if the recording is religious, research shows that this does not unduly bias the model to produce more religious language. According to the researchers, this is because it employs a more restrictive connectionist temporal classification strategy than large-scale language models (LLMs) and sequence-to-sequence models for speech recognition.
The team preprocessed the data by combining a highly efficient forced alignment approach capable of handling recordings of 20 minutes or longer with alignment models trained using data from over 100 different languages. To eliminate potentially distorted information, we repeated this procedure multiple times and used a cross-validation filtering step based on model accuracy. They integrated the alignment technique into his PyTorch and published the alignment model so that other scholars could use it to generate new speech datasets.
There is not enough information to train a traditional supervised speech recognition model on just 32 hours of data per language. The team utilized his wav2vec 2.0 to train an effective system, significantly reducing the amount of labeled data previously required. Specifically, we trained a self-supervised model on over 500,000 hours of his speech data, about five times more than previous efforts, using over 1,400 unique languages.
Researchers leveraged existing benchmark datasets such as FLEURS to evaluate the performance of models trained on large-scale multilingual speech data. His wav2vec 2.0 model with 1B parameters was used to train a multilingual speech recognizer in over 1,100 languages. Performance degrades slightly as the number of languages increases. From 61 to 1,107 languages, the misspelling rate only increases by about 0.4%, but the language coverage increases nearly 18 times.
Comparing Massively Multilingual Speech data with OpenAI’s Whisper, the researchers found that models trained on the former achieved half the word error rate. At the same time, the latter covers 11 times more languages than his. This indicates that this model can compete favorably with state-of-the-art speech recognition.
The team also trained language identification (LID) models for over 4,000 languages using their own datasets and publicly available datasets such as FLEURS and CommonVoice. We then tested it with the FLEURS LID challenge. The findings show that even with 40 times more languages supported, performance is still excellent. We have also developed a text-to-speech system that supports over 1,100 languages. Most existing text-to-speech algorithms are trained on single-speaker speech datasets.
The team envisions a world where one model can handle many speech tasks across all languages. They trained separate models for each task, such as language recognition, synthesis, and discrimination, but in the future, a single model could handle all these features and more, and any I believe it will improve the performance of the region.
Please check paper, blog, and Github link. don’t forget to join 22,000+ ML SubReddits, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more. If you have any questions regarding the article above or missed something, feel free to email me. Asif@marktechpost.com
🚀 Check out 100’s of AI Tools at the AI Tools Club
Tanushree Shenwai is a consulting intern at MarktechPost. She is currently pursuing her bachelor’s degree at the Indian Institute of Technology (IIT), Bhubaneswar. She is a data her science enthusiast and has a keen interest in the range of applications of artificial intelligence in various fields. She is passionate about exploring new advances in technology and its practical applications.
