Google Research Announces SPAE: AutoEncoder for Multimodal Generation Using Frozen Large Language Models (LLMs)

AI and ML Jobs


https://arxiv.org/abs/2306.17842

Large-scale language models (LLMs) are rapidly gaining great popularity due to their extraordinary capabilities in natural language processing and natural language understanding. This recent development in the field of artificial intelligence has revolutionized the way humans and computers interact with each other. A recent model developed by OpenAI is the well-known ChatGPT. Based on GPT’s Transformer architecture, this model is famous for mimicking humans for realistic conversations, doing everything from question answering and content generation to code completion, machine translation, and text summarization.

LLM excels at capturing deep conceptual knowledge about the world through lexical embeddings. However, researchers continue to strive to enable frozen LLMs to complete visual modality tasks when given appropriate visual representations as input. The researchers propose to utilize a vector quantizer that maps the image onto the frozen LLM’s token space. This translates the image into a language that his LLM understands, allowing for conditional image understanding and execution utilizing the LLM’s generative powers. A generative task that does not need to be trained on image-text pairs.

To address this issue and facilitate this crossmodal task, a team of researchers at Google Research and Carnegie Mellon University developed the Semantic Pyramid AutoEncoder, an autoencoder for multimodal generation using large-scale frozen language models. (SPAE) was introduced. SPAE produces lexical word sequences that convey rich semantics while preserving detail for signal reconstruction. In SPAE, the team combined an autoencoder architecture with a hierarchical pyramid structure. Unlike previous approaches, SPAE encodes images into interpretable discrete latent spaces, or words.

🚀 Check out 100’s of AI Tools at the AI ​​Tools Club

The pyramidal representation of SPAE tokens has multiple scales, with the lowest layer of the pyramid prioritizing appearance representations that capture details for image reconstruction, and the upper layers of the pyramid semantically capturing central concepts. includes. The system uses fewer tokens for tasks that require knowledge and more tokens for jobs that require generation, allowing the token length to be dynamically adjusted to accommodate different tasks. Adjustable. This model was trained independently without backpropagating through the language model.

To evaluate the effectiveness of SPAE, the team conducted experiments on image understanding tasks such as image classification, image captioning, and visual question answering. The results demonstrate how well LLM can handle visual modalities and several excellent applications such as content generation, design support, and interactive storytelling. The researchers also used in-context denoising techniques to illustrate the image generation capabilities of LLM.

The team summarizes their contributions as follows:

  1. This work provides an excellent method for directly generating visual content using in-context learning with a frozen language model trained only on language tokens.
  1. A Semantic Pyramid AutoEncoder (SPAE) has been proposed to generate interpretable representations of semantic concepts and details. The multilingual tokens generated by the tokenizer can be customized in length, allowing greater flexibility and adaptability in capturing and conveying subtleties of visual information.
  1. Progressive prompting techniques have also been introduced, enabling seamless integration of verbal and visual modalities, enabling the generation of comprehensive and consistent crossmodal sequences with improved quality and accuracy.
  1. This approach outperforms state-of-the-art few-shot image classification accuracy by an absolute margin of 25% under identical in-context conditions.

In conclusion, SPAE is an important advance that bridges the gap between language models and visual understanding. This demonstrates the remarkable potential of LLM in processing cross-modal tasks.


Please check paper.don’t forget to join 26,000+ ML SubReddit, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more. If you have any questions regarding the article above or missed something, feel free to email me. Asif@marktechpost.com

🚀 Check out 800+ AI tools in the AI ​​Tools Club

Tanya Malhotra is a final year student at the University of Petroleum and Energy Research, Dehradun, graduating with a Bachelor of Science in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
A data science enthusiast with good analytical and critical thinking, she has a keen interest in learning new skills, leading groups, and managing work in an organized manner.

🚀 Transform Selfies into AI Generated Headshots: Try the #1 AI Headshot Generator Today



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *