The transformer concept is widely accepted and applied in several areas of research and business. The most serious flaw of this model is the quadratic complexity of attention manipulation. This makes it difficult to apply large models to longer inputs. This study shows how a single Nvidia GTX 1080Ti GPU can process sequences of over a million tokens utilizing a straightforward token-based memory scheme combined with pre-trained transform models such as BERT. is shown.
A first step in making recurrent memory (RMT) generalizable to problems with unknown features such as language modeling is the study of synthetic tasks. Since this design became popular, much research has been done on the problem of long inputs in Transformers. This research shows that a significant amount of memory may be required only when using Transformer to analyze long texts. Repetitive strategies and memorization can transform quadratic complexity into linear complexity. Additionally, models trained with sufficiently large inputs may generalize to readers with longer orders of magnitude. They plan to modify the recurrent memory method in future work to increase the effective context size of the most frequently used Transformers.
Researchers from DeepPavlov, the Artificial Intelligence Research Institute, and the London Institute for Mathematical Sciences will make the following contributions:
1. Token-based memory storage and segment-level recursion (RMT) using recursive memory are added to BERT to improve the existing system.
2. They show that memory-enhanced BERTs can be trained to handle jobs with sequences up to 7 times longer than the intended input length of 512 tokens.
3. They found that the trained RMT can be effectively estimated for tasks of various durations, including tasks that require linear scaling of computation and tasks that exceed 1 million tokens.
4. Using attention pattern analysis, we discovered the memory processes that RMT uses to successfully process very long sequences.
The use of recursive memory in BERT, one of the most successful Transformer-based models in natural language processing, is presented by the authors as a conclusion. We effectively extended the model’s effective context length to an unprecedented 2 million tokens while maintaining excellent memory retrieval accuracy using the Recurrent Memory Transformer architecture. Their approach uses recursion to allow information flow between segments of the input sequence, allowing local and global information storage and processing. Their tests improve the handling of long-term dependencies in tasks involving natural language creation and understanding, and the effectiveness of a method with great potential for enabling large-scale contextual processing in memory-intensive applications. is shown.
check out paperdon’t forget to join 20,000+ ML SubReddit, cacophony channeland email newsletterWe share the latest AI research news, cool AI projects, and more. If you have any questions about the article above or missed something, feel free to email me. Asif@marktechpost.com
🚀 Check out 100 AI Tools in the AI Tools Club
Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing a Bachelor’s Degree in Data Science and Artificial Intelligence from the Indian Institute of Technology (IIT), Bhilai. He spends most of his time on projects aimed at harnessing the power of machine learning. His research interest is image processing and his passion is building solutions around it. He loves connecting with people and collaborating on interesting projects.
