Google's new TurboQuant algorithm speeds up AI memory 8x, cutting costs by 50% or more

TL;DR AI
2 min readKey summary
Google Research released the TurboQuant algorithm suite and published the associated research papers for free.
TurboQuant reduces KV memory use by 6x and speeds computation of attention logits by 8x.
Enterprises that adopt it could cut model costs by more than 50%, and it is available for enterprise use.
The underlying math (PolarQuant and Quantized Johnson‑Lindenstrauss) was documented in early 2025; the work stems from a multi‑year research arc starting in 2024 and was formally unveiled today.
Google frames these methods as plumbing for the Agentic AI era; the release times with ICLR 2026 and AISTATS 2026 and is believed to put pressure on memory‑provider stocks.



