Switch language한국어
Back to the list

GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia

TL;DR AI

Key summary

2 min read
  1. Z.ai launched GLM-5.3-Flash, an MIT-licensed open-weight multimodal model with 320B total parameters, 18B active parameters, and a 1M-token context window.

  2. The model reportedly approaches larger frontier systems on benchmarks while offering significantly lower token costs.

  3. Z.ai says it deployed inference at scale on Chinese AI chips rather than Nvidia hardware, challenging both model pricing and reliance on the CUDA ecosystem.

Read the original