TECH·May 30, 2026Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA | Hacker NewsHacker News
TECH·May 27, 2026Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM InferenceMarkTechPost
TECH·May 27, 2026Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team | Hacker NewsHacker News
HARDWARE·May 12, 2026AMD’s vLLM-ATOM Plugin Supercharges DeepSeek-R1, Kimi-K2, and gpt-oss-120B AI LLM Inference on Instinct MI350 and MI400 AcceleratorsWccftech
PAPER·May 11, 2026UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic SparsificationHugging Face Papers
PAPER·April 30, 2026Accelerating RL Post-Training Rollouts via System-Integrated Speculative DecodingHugging Face Papers
TECH·April 26, 2026A Coding Implementation on kvcached for Elastic KV Cache Memory, Bursty LLM Serving, and Multi-Model GPU SharingMarkTechPost