Qwen-Audio-3.1: Alibaba's Five-Model Audio Stack
Melvin Vivas · X post · 2026-09-23 · Open on X
Topics: Industry Trends & Job Market, LLM Fundamentals · Level: beginner
Summary
Alibaba Qwen released Qwen-Audio-3.1, which upgrades its speech recognition (ASR), text-to-speech (TTS) and Realtime models. It also adds two new models: TTS-Next for audio creation and ASR-Next for audio understanding. Together the five models cover understanding, generation, interaction and creation, and Qwen announced big price cuts.
Key points
- Qwen-Audio-3.1 upgrades the ASR, TTS and Realtime models.
- New TTS-Next model: for audio creation.
- New ASR-Next model: for audio understanding.
- Five models make up one audio stack covering understanding, generation, interaction and creation.
- Qwen announced big price cuts across the line-up (the quoted post is cut off before the details).
Resources mentioned
- Qwen (@Alibaba_Qwen) on X · person · x.com · free
Official X account of Alibaba's Qwen team, which posts model releases and demos.
Also in: Demo: Qwen3.8-27B Running Locally with Pi and llama.cpp (Melvin Vivas on X · notes), Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models (Melvin Vivas on X · notes), Audio-Visual Vibe Coding with Qwen 3.5 Omni: Spoken Specs to Web App (Melvin Vivas on X · notes) - Qwen TTS-Next · tool · x.com · paid
Alibaba's audio model family with upgraded ASR, TTS and Realtime models.
Try this
- Follow @Alibaba_Qwen for audio model releases and compare Qwen's ASR/TTS models for voice applications.
More in Industry Trends & Job Market
- ChatGPT Voice Update: Plugins, Model Switching and ChatGPT Work
- Gemini 3.8 Flash and Flash-Lite TTS: new text-to-speech models
- Space Bunny Alpha: stealth 1M-context flash model on OpenRouter
- DeepSWE results: GPT-6 Sol slightly below GPT-5.6 Sol, but cheaper
- OpenAI Releases GPT-6 Sol and GPT-6 Luna
- Claude Opus 5.5 Released: Fable 5.1-Level Performance at 40% Lower Cost