[REVISION HISTORY]
Zhipu AI release of GLM-5.3 and GLM-5.3-Flash models
Updated 12 times since CLSTR started tracking revisions of this situation.
What changed
2026-09-07 05:47 UTC → 2026-09-18 05:44 UTC ·
added
removed
Zhipu AI has released GLM-5.3 and its optimized sibling, GLM-5.3-Flash, as open-weight models. GLM-5.3 utilizes a 744-billion-parameter Mixture of Experts architecture optimized for programming and cybersecurity, demonstrating state-of-the-art performance on benchmarks like Terminal Bench 3.0 and achieving an 84.5% success rate in CyberGym. While it excels at vulnerability identification, it has trailed Western competitors in complex exploit chain execution. The GLM-5.3-Flash variant features a 320-billion-parameter Mixture of Experts architecture with 18 billion active parameters per token. It is a multimodal model capable of natively processing text, images, and video, utilizing a hybrid architecture of linear and sparse attention to manage its 1-million-token context window. Before its official release, the model was distributed anonymously as ‘ox-alpha’ on platforms such as OpenRouter, where it became the most used model within a single week, reportedly exceeding 50 trillion tokens in traffic. On the Artificial Analysis Intelligence Index, GLM-5.3-Flash secured a top-ten position, performing similarly to Anthropic’s Claude Opus 4.8. A significant technical milestone is the model’s reliance on domestic Chinese compute infrastructure. Its online inference traffic is serviced by an array of approximately 100,000 domestically manufactured AI chips, such as those from Cambricon; analysts suggest the hardware may belong to Huawei’s Ascend series. Following the release, the US startup Abliteration.ai developed a modified version titled ‘abliterated-model-large-v2’. This ‘abliteration’ process modifies model weights to suppress internal activation patterns associated with refusal mechanisms, aiming to create a version less likely to decline sensitive or security-related queries while maintaining coding and cyber capabilities. Zhipu AI has since expanded the lineup with the launch of GLM-5.3-FlashX. This high-speed version is designed for enterprises and developers, offering a maximum inference speed of up to 200 tokens per second. While it maintains the 320-billion-parameter architecture and 1-million-token context window of the standard Flash model, it is priced higher.
Versions
- 2026-09-18 05:44 UTC Zhipu AI release of GLM-5.3 and GLM-5.3-Flash models
- 2026-09-07 05:47 UTC Zhipu AI release of GLM-5.3 and GLM-5.3-Flash models
- 2026-08-31 07:42 UTC Zhipu AI release of GLM-5.3 and GLM-5.3-Flash models
- 2026-08-29 03:27 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-29 02:55 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-27 23:34 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-27 15:32 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-26 21:31 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-23 07:32 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-22 10:41 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-18 23:40 UTC Zhipu AI release and autonomous capabilities of GLM-5.3
- 2026-08-17 01:39 UTC Zhipu AI release of GLM-5.3 model
- 2026-08-16 15:16 UTC Zhipu AI release of GLM-5. 3 model
Only revisions since CLSTR began indexing content versions appear here. Select a version to see what changed compared to the one before it.