Paper·May 21, 2026OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and BeyondHugging Face Papers
Paper·May 14, 2026Orthus: Memory-Efficient Parallel Token Generation via Dual-View DiffusionHugging Face Papers
Paper·May 11, 2026SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree DraftingHugging Face Papers