TECH·4 days agoRelay-style orchestration revitalizes national computing power, and PD separation finally breaks through: latency cut in half, costs down nearly 40%!量子位
TECH·6 days agoChina Surpasses the US, Accounting for Two-Thirds of Global LLM Calls, as Xiaomi MiMo-V2.5 Ranks First Worldwide in Weekly and Monthly Token UsagePandaily
PAPER·6 days agoBeyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator SchedulingarXiv
HARDWARE·July 26, 2026AI enthusiast adds Nvidia Tesla V100 as loud as a lawnmower to gaming PC for $266 — 32GB VRAM rig can run 27 billion parameter model at 32 tokens per secondTom's Hardware
PAPER·June 2, 2026Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative DecodingHugging Face Papers
TECH·May 30, 2026Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA | Hacker NewsHacker News
PAPER·May 30, 2026CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLMHugging Face Papers
TECH·May 29, 2026Real-time LLM Inference on Standard GPUs: 3k tokens/s per request | Hacker NewsHacker News
TECH·May 27, 2026Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM InferenceMarkTechPost