Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

TL;DR AI
2 min readKey summary
Researchers at UC Berkeley and UT Austin introduced FreeToken, an edge-native serving engine for running MoE LLMs on a single personal machine.
FreeToken adapts inference to available GPU, CPU, memory, and PCIe bandwidth, reportedly supporting models from 35B and 284B up to 753B GLM-5.2.
Released under Apache-2.0, it offers a CLI, PyPI package, desktop app support, and an OpenAI-compatible endpoint.
It could help developers and small teams cut cloud inference costs while keeping sensitive workloads on-device.



