Switch language한국어
Back to the list

Alibaba Qwen Team Releases Qwen3.5 Omni: A Native Multimodal Model for Text, Audio, Video, and Realtime Interaction

TL;DR AI

Key summary

2 min read
  1. Qwen3.5-Omni uses a Thinker-Talker architecture.

  2. It includes a native Audio Transformer (AuT) encoder pre-trained on more than 100 million hours of audio-visual data.

  3. Both the Thinker and the Talker leverage Hybrid-Attention Mixture-of-Experts (MoE).

  4. The architecture supports a 256k long-context input and can ingest over 10 hours of continuous audio.

  5. The article also mentions "over 400 seconds" related to audio/video, but that line is incomplete in the provided content.

Read the original