768GB of cheap Intel Optane DIMM memory sticks used to run 1-trillion-parameter LLM on a system with a single GPU — local Kimi K2.5 install achieved roughly 4 tokens per second

TL;DR AI
2 min readKey summary
A Redditor ran the 1-trillion-parameter Kimi K2.5 locally with llama.cpp on a Xeon workstation.
The setup paired 768GB of used Intel Optane persistent memory, 192GB of DDR4, and a 12GB RTX 3060.
The system reached roughly 4 tokens per second, showing surprisingly usable inference on consumer and second-hand hardware.
The case highlights how low-cost Optane can bridge the gap between DRAM and SSDs for very large models.
