TECH·May 27, 2026Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM InferenceMarkTechPost
TECH·May 27, 2026Millions of AI agents are at risk due to a vulnerability in the open-source package Starlette, which is downloaded more than 300 million times a weekGIGAZINE
TECH·May 27, 2026Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRLHugging Face Blog
TECH·May 27, 2026Millions of AI agents imperiled by critical vulnerability in open source packageArs Technica
TECH·May 27, 2026Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team | Hacker NewsHacker News
PAPER·May 26, 2026ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM InferencearXiv
PAPER·May 22, 2026KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM ServingHugging Face Papers
PAPER·May 20, 2026OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache QuantizationHugging Face Papers
TECH·May 13, 2026MiniCPM-V 4.6: Tsinghua Spinoff Open-Sources a 1.3B Multimodal Model That Runs on a Single RTX 4090Pandaily