< Back to all clusters
[TECHNOLOGY] · China · 9 sources

started · updated

Z. ai releases GLM-5. 3 open-weight models

Z. ai has released GLM-5. 3 and its optimized sibling, GLM-5. 3-Flash, as open-weight models. The GLM-5. 3-Flash model utilizes a Mixture of Experts (MoE) architecture with 320 billion total parameters, of which 18 billion are active per token. It is designed for high-efficiency coding and agentic tasks, featuring a 1-million-token context window and native multimodality for text, images, and video.

A notable technical aspect of the release is its hardware compatibility; the model was tested and successfully run on approximately 100,000 Chinese AI chips, such as those from Cambricon, without relying on Nvidia GPUs. This demonstrates the ability of domestic Chinese hardware to support large-scale, frontier-level AI inference.

Performance benchmarks indicate that GLM-5. 3-Flash achieves competitive results, with an Artificial Analysis Intelligence Index score of 57, placing it on par with Anthropic’s Claude Opus 4.8. While the Flash variant shows slightly lower reliability in certain reasoning tasks compared to the full GLM-5. 3 model, it offers significant economic advantages, reducing rollout costs by approximately 17 times and providing higher throughput for cost-sensitive workloads.

Entities

Cambricon · Claude Opus 4. 8 · GLM-5. 3 · GLM-5. 3-Flash · OpenRouter · SenseTime · Z. ai