< Back to all clusters
[TECHNOLOGY] · 2 sources

started · updated

Alibaba releases Qwen3. 8-Flash open-weight multimodal model

Alibaba has released Qwen3. 8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model designed for high-volume applications, coding, and tool-driven workflows. The model features a 125B-parameter main architecture with 6B parameters activated per token, supplemented by 51B N-gram embeddings. It supports a native context of 262K tokens, extendable to 1 million tokens.

In the landscape of late 2026 local AI development, Qwen3. 8-Flash is noted for its efficiency on consumer hardware. The model is available via API on Alibaba’s Model Studio and Qwen Cloud, and its weights can be downloaded from Hugging Face and ModelScope. It is also integrated into the QwenWork platform, where it aims to reduce token consumption by 75% and increase generation speeds.

Entities

Alibaba · Hugging Face · ModelScope