< Back to all clusters
[TECHNOLOGY] · China · 5 sources

started · updated

Alibaba's Qwen Audio 3.0 Beats OpenAI in Speech Benchmark and Launches ASR-Flash Model

Alibaba released the Qwen-Audio-3.0-ASR-Flash speech‑recognition model, tuned for industry‑specific terminology and achieving a 95.36% accuracy rate in medical tests and a 1.7% typo rate on the Artificial Analysis platform. In a separate speech‑to‑speech benchmark, Qwen Audio 3.0 Real‑time Plus recorded a 99.2% reasoning score and 98.4% conversational dynamics, surpassing OpenAI’s comparable models, though its latency (about four seconds) was longer than OpenAI’s 1.14 seconds.

The Qwen family supports full‑duplex dialogue, function calling and custom voice cloning via Alibaba Cloud’s DashScope API, and is offered in three variants: ASR‑Flash for fast transcription, ASR‑Filetrans for offline files, and ASR‑Streaming for real‑time use. The technology is already being applied to meeting minutes, live subtitles, educational recordings and intelligent customer‑service bots.

Entities

Alibaba Group · OpenAI · Qwen Audio 3.0