started · updated
Zhipu AI launches GLM-5.3-Flash model powered by domestic chips
Zhipu AI has released GLM-5.3-Flash, a new 320B parameter multimodal large language model. The model utilizes a hybrid architecture combining linear attention for local dependencies and sparse attention for broad context, aiming to reduce costs for processing long-context windows.
Notably, the model's online inference traffic is reportedly serviced entirely by an array of 100,000 domestically manufactured Chinese AI chips. While Zhipu AI has not officially named the supplier, analysts suggest the hardware may belong to Huawei’s Ascend series. The deployment indicates a growing trend among Chinese AI developers to rely on domestic compute infrastructure for flagship workloads.
Prior to its official release, the model was distributed anonymously as ‘ox-alpha’ on platforms like OpenRouter and OpenCode, where it became the most used model within a single week. On the Artificial Analysis AAII leaderboard, GLM-5.3-Flash secured the tenth position, outperforming DeepSeek’s V4 Pro Max model.