Alibaba Cloud releases the “open” QWEN3-OMNI.

Machine Learning


Chinese tech giant Alibaba Cloud has announced a new entry for Qwen or Tongyu Qianwen, a family of large language models: Qwen3-Omni.

“We are releasing QWEN3-OMNI, a native end-to-end multilingual Omni-Modal Foundation model,” says Xiong Wang of Alibaba Cloud about the company's latest launch in a large-scale language model field. “It is designed to provide real-time streaming responses in both text and natural audio while processing a variety of inputs, including text, images, audio and video.”

QWEN3-OMNI promises “open” LLM with low latency suitable for speech-based interactions and true multimodal operation. (📹: Alibaba Cloud)

Large-scale language models are the technologies that support the current “artificial intelligence” boom. Usually, you will be ingested a statistical model that distills copyright, data, and data and distills it into a “token.” Then, it will be more tokens before responding with the most statistically limitless continuity token. When everything goes well, the answer shape coincides with reality. Otherwise, the response is only an answer with a divorced appearance from reality.

“QWEN3-OMNI employs a thinker and talker architecture,” the company's LLM development team said of the new model. “Though thinkers are in charge of text generation, while talkers focus on generating streaming speech tokens by receiving high-level representations directly from thinkers. To achieve ultra-low latency streaming, talkers automatically cover and predict multi-codebook sequences. They enable streaming generation per frame.”

However, at the same time, the ability to “think” and “speak” does not make QWEN3-OMNI interesting. Rather, it is the promise of a true multimodal model. A single model that can handle text, video, audio, and image input and output in one place. “By mixing unimodal data and cross-modal data in the early stages of pre-text registration, it claims that there is no performance degradation inherent to all modalities, namely modality, while “significantly enhancing cross-modal functionality.”

Company's Unique technical reportsHowever, pour a little cool water into this latter claim: QWEN3-OMNI shows strong performance on all media types, but by modern LLMS standards, performance during text processing is significantly weaker than previous QWEN3 intruct models.

Other features of the model include 119 languages ​​in text mode, 19 languages ​​for speech recognition, 10, 211ms low audio only latency, 507ms audio video latency support, support for up to 30 minutes of audio input, support for “tool-calling” – the ability to run “AI assistant” in order to run “AI assistant” in the order of being responsible for “handling AI assistants”. Details on how to do it yourself.

Alibaba Cloud has released three tailored variants of the model – QWEN3-OMNI-30B-A3B-INTRUCT, QWEN3-OMNI-30B-A3B-Thinking, and QWEN3-OMNI-30B-A3B-Captioner – github, Hugging my faceand ModelScope Under the generous Apache 2.0 license. However, as usual in the LLMS field, these are not really “open source” models, as they don't provide everything you need to build a model from scratch. Demos are also available Hugging my face.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *