GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia

TL;DR AI
2 min readKey summary
Z.ai launched GLM-5.3-Flash, an MIT-licensed open-weight multimodal model with 320B total parameters, 18B active parameters, and a 1M-token context window.
The model reportedly approaches larger frontier systems on benchmarks while offering significantly lower token costs.
Z.ai says it deployed inference at scale on Chinese AI chips rather than Nvidia hardware, challenging both model pricing and reliance on the CUDA ecosystem.



