PAPER·May 28, 2026The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World IntelligenceHugging Face Papers
PAPER·May 27, 2026Negligible in Size, Significant in Effect: On Scale Vectors in Large Language ModelsHugging Face Papers
PAPER·May 26, 2026Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEsHugging Face Papers
PAPER·May 22, 2026Semantically Structured Mixture-of-Experts for Compositional Robotic ManipulationarXiv
PAPER·May 22, 2026TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert OffloadHugging Face Papers
PAPER·May 18, 2026HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-ExpertsHugging Face Papers
PAPER·May 15, 2026BEAM: Binary Expert Activation Masking for Dynamic Routing in MoEHugging Face Papers
TECH·May 14, 2026Nous Research Releases Token Superposition Training to Speed Up LLM Pre-Training by Up to 2.5x Across 270M to 10B Parameter ModelsMarkTechPost
TECH·May 13, 2026Meet AntAngelMed: A 103B-Parameter Open-Source Medical Language Model Built on a 1/32 Activation-Ratio MoE ArchitectureMarkTechPost