Alibaba Expands Qwen Into Speech And Mobile Agents As Qwen 4 And Custom Silicon Signal Broader Ambitions

Alibaba’s Qwen team has released two updates to its product lineup: Qwen Audio 3.1, a rebuilt speech stack, and Qwen Intelligence, a suite of mobile agents. Together they illustrate how the company is converting its model momentum into deployable products across modalities.
Qwen Audio 3.1 comprises five models organized into two groups: an upgraded core lineup covering understanding, generation, and interaction, and two new models aimed at creation and deeper comprehension. On the synthesis side, the flagship TTS model ranks first on the independent Artificial Analysis Text to Speech Leaderboard. It supports 16 languages and 20 Chinese dialect regions, produces up to three minutes of continuous speech in a single pass, and accepts free style natural language instructions controlling emotion, pace, timbre, and accent, alongside 86 fine grained inline tags for phrase level effects such as pauses, breathing, laughter, and sighs.
According to the technical report, a 12.5 Hz low frame rate speech tokenizer reduces decoding cost, while a five stage training pipeline coordinates the language and flow models for content consistency, voice similarity, and robustness; the system generates usable speech even from noisy, reverberant, or degraded reference audio. A companion model, TTS Next, combines a language model with a diffusion framework to generate voice, sound effects, and background audio in one pass, targeting audiobooks, podcasts, games, and advertising.
On the understanding side, the upgraded ASR adds stronger multilingual and dialect recognition with native transcript polishing that removes fillers and repetitions automatically. ASR Next extends this to multi speaker recognition with speaker labels, timestamps, and aligned transcripts, and can interpret emotions, ambient noise, and machine sounds, enabling sound captioning, event localization, and audio question answering.
The Realtime model supports simultaneous speaking and listening with interruption at any moment, and adjusts its pace and tone when it detects a low mood in the caller. Alibaba paired the launch with substantial price reductions: TTS costs dropped roughly 70%, Realtime around 85%, and ASR by up to 95%.
Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.Five models, one complete audio stack: understanding, generation, interaction & creation. Plus big price cuts across the… pic.twitter.com/Qcbvcw0q9C
Qwen Intelligence targets a different layer entirely: autonomous task execution on mobile devices. Its Planner agent decomposes and orchestrates complex tasks, ranking first on the company’s MobilePA Bench, including its business and memory variants. The Mobile Use agent executes actions primarily through APIs, falling back to graphical interface control when needed, scoring 82.1 on MobileWorld, 92.2 on the real device MobileWorld Real benchmark, and 97.2 on AndroidDaily, with a reported 90 percent success rate on complete end to end tasks.
The Creative agent generates a usable image from a single sentence in about three seconds, roughly twice as fast as leading alternatives, Alibaba claims. Notably, the company also opened its benchmark suite, including MobilePA Bench, MobileWorld, MobileWorld Real, and MobileWorld Safety, to the research community, providing external reference points for evaluating mobile agents.
Introducing Qwen Intelligence, bringing personal intelligence within everyone's reach. 
It launches with three SOTA agents:
– Mobile Planner Agent: plans, decomposes & orchestrates complex tasks. #1 on MobilePA-Bench, MobilePA-Bench Business & Memory.– Mobile-Use Agent:… pic.twitter.com/OgCLpxDJgj
These releases arrive in the wake of Alibaba’s Apsara Conference in Hangzhou on September 22, where the Qwen 4 family was announced in four tiers: Max as the flagship positioned against top rival models, Flash for low latency and high volume workloads, Plus as a balanced multimodal tier, and a 27B open weights variant for local use. None of the four has public specifications, pricing, or a launch date, making the announcement a statement of direction rather than a shipping product.
The conference’s other announcements clarify that direction. Chief Executive Eddie Wu stated that the Qwen team plans to train models of five to ten trillion parameters, a target Alibaba’s own statements assign to Qwen 4.5 and Qwen 5, for tackling more complex, longer horizon tasks. To support models of that scale, Alibaba’s T Head unit unveiled the Zhenwu V900 accelerator, which promises three times the performance of the M890 predecessor, carries 216 GB of memory, and scales to clusters of up to 500,000 chips, with mass production scheduled for the first quarter of 2027.
Wu set a target of surpassing 20 gigawatts of global cloud capacity by 2032 and described a three layer cloud architecture designed for agentic workloads. He also noted that demand for AI is outpacing supply, with AI supernodes coming online at commercial scale this quarter. Alibaba’s Hong Kong shares rose 5.1 percent on the day. The domestic chip push unfolds against tightening US export curbs, while analysts point to August’s open sourced Qwen3.8 Flash Next, with its sparse attention and n gram embedding techniques, as the closest available preview of the next generation architecture, though Alibaba has not confirmed the connection.
A Broader Trajectory: From Open Models to an Integrated StackThe recent releases cap an accelerating trajectory. Since its debut in 2023, Qwen has become the most widely used open weights model family, with more than 300 million cumulative downloads and over 100,000 derivative models on Hugging Face, according to Alibaba.
The current flagship, Qwen3.8 Max, released in August with open weights, contains 2.4 trillion parameters and handles one million tokens of context, while September’s Qwen3.8 Omni Flash extends native processing across text, image, audio, and video. The licensing structure has evolved in parallel: the flagship carries a clause requiring companies generating over $50 million in annual revenue to share that revenue with Alibaba, while smaller distilled models ship under permissive Apache 2.0 licenses.
Several open questions will shape how this strategy unfolds: whether independently verified benchmarks confirm the vendor reported results for agents and speech models, whether chip production timelines align with the model roadmap, and how the revenue sharing license is adopted in practice. With Qwen 4 specifications still undisclosed and the next generation already in training, Alibaba has laid out an unusually broad agenda spanning silicon, cloud infrastructure, models, and consumer agents. The coming quarters will show how these pieces converge in practice, and we will keep watching how Qwen develops.
The post Alibaba Expands Qwen Into Speech And Mobile Agents As Qwen 4 And Custom Silicon Signal Broader Ambitions appeared first on Metaverse Post.


Metaverse Post





