< Back to all clusters
[TECHNOLOGY] · China · 6 sources

started · updated

Zhipu AI launches GLM-5.3-FlashX with 200 tokens/s speed

Zhipu AI has launched GLM-5.3-FlashX, a high-speed version of its large language model designed to provide faster and smoother experiences for enterprises and developers. The model boasts a maximum inference speed of up to 200 tokens per second, supported by an infrastructure utilizing 100,000 domestic chips.

GLM-5.3-FlashX is built upon the GLM-5.3-Flash base model, which features 320 billion total parameters and 18 billion active parameters per inference. It supports a context window of up to 1 million tokens. The release includes the availability of an API and an experience center for testing model responses.

While the FlashX version offers enhanced speed, it is priced higher than the standard GLM-5.3-Flash. The model supports multiple input modalities, including text, images, video, and files.

Entities

Zhipu AI

Sources

7项评测领先开源模型 宇树科技开源具身基座模型让人形机器人“理解世界”_天极网 [news.yesky.com]