< Back to all clusters
[TECHNOLOGY] · China · 24 sources

started · updated

Z. ai and Alibaba release new efficient open-weight AI models

Chinese AI firm Z. ai has released GLM-5. 3-Flash, an open-weight multimodal model previously known under the codename Ox Alpha. The model utilizes a Mixture-of-Experts (MoE) architecture with 320 billion total parameters, activating approximately 18 billion per token. It supports text, image, and video inputs with a context window of up to 1 million tokens. Z. ai claims the model is significantly more cost-efficient than its predecessor and noted that its initial testing phase on platforms like OpenRouter was supported by Chinese AI chip clusters.

Simultaneously, Alibaba has introduced Qwen3. 8-Flash-Next, an open-weight model serving as a preview of its upcoming Qwen4 architecture. This 125-billion-parameter model activates 6 billion parameters per token and features 51 billion N-gram embeddings. It incorporates architectural innovations such as Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to optimize long-context processing. The model is designed for high-efficiency tasks including agentic coding and document analysis, with support for NVIDIA GB300 NVL72 platforms.

Entities

Alibaba · Anthropic · Hugging Face · ModelScope · Nvidia · OpenRouter · Qwen · Z. ai

Claims

What the coverage asserts, and how many sources carry each claim.

Sources

16 days ago