started · updated
Zhipu AI launches GLM-5.3-FlashX with 200 tokens/s speed
Zhipu AI has launched GLM-5.3-FlashX, a high-speed version of its large language model designed to provide faster and smoother experiences for enterprises and developers. The model boasts a maximum inference speed of up to 200 tokens per second, supported by an infrastructure utilizing 100,000 domestic chips.
GLM-5.3-FlashX is built upon the GLM-5.3-Flash base model, which features 320 billion total parameters and 18 billion active parameters per inference. It supports a context window of up to 1 million tokens. The release includes the availability of an API and an experience center for testing model responses.
While the FlashX version offers enhanced speed, it is priced higher than the standard GLM-5.3-Flash. The model supports multiple input modalities, including text, images, video, and files.