Introducing VampNet: A Masked Acoustic Token Modeling Approach to Music Synthesis, Compression, Restoration, and Variation

AI and ML Jobs


https://arxiv.org/abs/2307.04686

Recently, the development of discrete acoustic token modeling has greatly improved speech and music autoregressive production. A non-autoregressive parallel iterative decoding method was devised for efficient image creation. Such infill jobs are more suitable for parallel iterative decoding than autoregressive approaches because they require conditioning on both past and future sequence components. This work utilizes acoustic token modeling and joint iterative decoding to music-to-speech synthesis. To their knowledge, they are the first to use parallel iterative decoding for neural audio music synthesis.

They use token-based prompts to adapt their model, known as VampNet, to a wide range of applications. Through deliberately concealed musical token sequences, they demonstrate his ability to direct the creation of VampNet and fill in the blanks. The results of this process range from high-quality audio compression methods to music that closely resembles the original input music in style, genre, beats, and instrumentation, but with some altered timbral and rhythmic nuances. you will get something. Their method allows prompts to be placed anywhere, unlike autoregressive music models that can only perform musical continuations by utilizing the prefix audio as a prompt and having the model generate the music that follows it.

Figure 1: Overview of VampNet. First, split the audio into a series of individual tokens using an audio tokenizer. Tokens are first masked before being sent to a masked generative model that uses an efficient iterative parallel decoding sampling technique at two levels to predict the value of the masked token. The output is then decoded to audio.

They explore a variety of prompt designs, including periodic, compressed, and music-inspired (such as masking on beats). They found that the model performed brilliantly when told to create loops and variations. Hence the name VampNet. They provide a code for download and urge people to check out the audio samples. Researchers at Descript Inc. and Northwestern University introduced his VampNet, a method for generating music using masked acoustic token modeling. VampNet is bi-directional, so the input audio file may prompt in different ways. VampNet is a great tool for creating musical variations as it works continuously between music compression and production through different prompting approaches.

🚀 Check out 100’s of AI Tools at the AI ​​Tools Club

A musician can use VampNet to record a short loop, feed it into the system, and let VampNet devise a musical variation based on the idea each time the loop region is repeated. They plan to study his VampNet, facilitating potential approaches to interactive musical co-production in further research and to the representational learning capabilities of masked acoustic token modeling.


Please check paper. All credit for this research goes to the researchers of this project.Also, don’t forget to participate 26,000+ ML SubReddit, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more.

🚀 Check out 800+ AI Tools in the AI ​​Tools Club

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his Bachelor of Science in Data Science and Artificial Intelligence from the Indian Institute of Technology (IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is in image processing and he is passionate about building solutions around it. He loves connecting with people and collaborating on interesting projects.

🚀 Build high-quality training datasets, solve NLP machine learning challenges, and develop powerful ML applications with Kili Technology



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *