TECH·June 3, 2026Tether Brings Google TurboQuant to Everyday Devices, Giving Local AI Data Center-Sized MemoryTekedia
TECH·May 26, 2026Together AI Open-Sources OSCAR: An Attention-Aware 2-Bit KV Cache Quantization System for Long-Context LLM ServingMarkTechPost
CODING·May 23, 2026We Replaced Our RAG Pipeline With Persistent KV Cache. Here's What We Found.Dev.to
PAPER·May 22, 2026Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training StepsHugging Face Papers
PAPER·May 21, 2026OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under Optimal Squared Error QuantizationHugging Face Papers
PAPER·May 21, 2026OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and BeyondHugging Face Papers
PAPER·May 20, 2026OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under Optimal Squared Error QuantizationarXiv
PAPER·May 19, 2026CompactAttention: Accelerating Chunked Prefill with Block-Union KV SelectionHugging Face Papers