started · updated
Z. ai and Alibaba release new efficient open-weight AI models
Chinese AI firm Z. ai has released GLM-5. 3-Flash, an open-weight multimodal model previously known under the codename Ox Alpha. The model utilizes a Mixture-of-Experts (MoE) architecture with 320 billion total parameters, activating approximately 18 billion per token. It supports text, image, and video inputs with a context window of up to 1 million tokens. Z. ai claims the model is significantly more cost-efficient than its predecessor and noted that its initial testing phase on platforms like OpenRouter was supported by Chinese AI chip clusters.
Simultaneously, Alibaba has introduced Qwen3. 8-Flash-Next, an open-weight model serving as a preview of its upcoming Qwen4 architecture. This 125-billion-parameter model activates 6 billion parameters per token and features 51 billion N-gram embeddings. It incorporates architectural innovations such as Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to optimize long-context processing. The model is designed for high-efficiency tasks including agentic coding and document analysis, with support for NVIDIA GB300 NVL72 platforms.
Entities
Alibaba · Anthropic · Hugging Face · ModelScope · Nvidia · OpenRouter · Qwen · Z. ai
Claims
What the coverage asserts, and how many sources carry each claim.
- [● 12 SOURCES] Z. ai released GLM-5. 3-Flash, an open-weight multimodal model previously known by the codename Ox Alpha. www.prodiris.fr · www.sbctv.gr · www.it-boltwise.de · siliconangle.com · dev.to · +6 more
- [● 7 SOURCES] The model activates approximately 6 billion parameters per token. bitcoinethereumnews.com · www.ule.co.id · thetechnologyexpress.com · thenextweb.com · decrypt.co · +2 more
- [● 10 SOURCES] The model has a native context window of 262,144 tokens, which can be extended to 1 million tokens via the YaRN framework. bitcoinethereumnews.com · www.ule.co.id · thetechnologyexpress.com · developer.nvidia.com · gigazine.net · +5 more
- [● 10 SOURCES] GLM-5. 3-Flash features 320 billion total parameters with 18 billion activated per prompt. www.prodiris.fr · siliconangle.com · pivot.uz · www.sbctv.gr · www.itmedia.co.jp · +4 more
- [● 6 SOURCES] Qwen3. 8-Flash-Next features 125 billion parameters and 51 billion N-gram embeddings. bitcoinethereumnews.com · www.ule.co.id · thetechnologyexpress.com · thenextweb.com · gigazine.net · +1 more
- [● 7 SOURCES] GLM-5. 3-Flash is reported to be ten times more cost-efficient than its predecessor. www.prodiris.fr · siliconangle.com · pivot.uz · 4sysops.com · the-decoder.de · +1 more
- [● 3 SOURCES] The Qwen Sparse Attention (QSA) kernel reportedly provides up to 7.6x speedup in prefill and 4.9x in decoding compared to traditional mechanisms. bitcoinethereumnews.com · www.ule.co.id · developer.nvidia.com
- [● 7 SOURCES] Alibaba released Qwen3. 8-Flash-Next as an open-weight preview of its upcoming Qwen4 architecture. bitcoinethereumnews.com · thetechnologyexpress.com · www.ule.co.id · decrypt.co · thenextweb.com · +2 more
- [● 2 SOURCES] Qwen3. 8-Flash-Next uses Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA). bitcoinethereumnews.com · developer.nvidia.com
- [● 3 SOURCES] Z. ai claims the model was tested and run on Chinese AI chip clusters. www.itmedia.co.jp · the-decoder.de · abmedia.io