TECH·May 27, 2026Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team | Hacker NewsHacker News
TECH·May 16, 2026Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution | Hacker NewsHacker News
TECH·May 13, 2026Meet AntAngelMed: A 103B-Parameter Open-Source Medical Language Model Built on a 1/32 Activation-Ratio MoE ArchitectureMarkTechPost
TECH·May 7, 2026Google announces "multi-token prediction," a technology that uses a small AI to generate drafts and speed up large AIGIGAZINE
TECH·May 7, 2026Google's Gemma 4 AI models get a 3x speed boost by predicting future tokensArs Technica
TECH·May 6, 2026Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without Quality LossMarkTechPost
TECH·May 6, 2026Accelerating Gemma 4: faster inference with multi-token prediction draftersHacker News