Zhipu AI Launches GLM-5.1 High-Speed API: 400 Tokens/s Sets New Global Benchmark

TL;DR AI
2 min readKey summary
Zhipu AI has launched GLM-5.1-highspeed, a faster API version of its GLM-5.1 model.
The rollout is limited to selected enterprise customers and targets latency-sensitive use cases like content generation, coding support, and customer interaction.
Zhipu says the service can reach 400 tokens per second, setting a new benchmark for LLM inference speed.
The move underscores rapid progress in real-time model serving and raises competitive pressure across top AI providers.
