Switch language한국어
Back to the list

768GB of cheap Intel Optane DIMM memory sticks used to run 1-trillion-parameter LLM on a system with a single GPU — local Kimi K2.5 install achieved roughly 4 tokens per second

TL;DR AI

Key summary

2 min read
  1. A Redditor ran the 1-trillion-parameter Kimi K2.5 locally with llama.cpp on a Xeon workstation.

  2. The setup paired 768GB of used Intel Optane persistent memory, 192GB of DDR4, and a 12GB RTX 3060.

  3. The system reached roughly 4 tokens per second, showing surprisingly usable inference on consumer and second-hand hardware.

  4. The case highlights how low-cost Optane can bridge the gap between DRAM and SSDs for very large models.

Read the original