PAPER·5 days agoThe Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic DistillationHugging Face Papers
TECH·May 30, 2026Making AI chatbots helpful weakens their ability to simulate human behavior, large-scale study findsTHE DECODER
PAPER·May 29, 2026RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable DomainsHugging Face Papers
PAPER·May 28, 2026Guiding LLM Post-training Data Engineering with Model Internals from Sparse AutoencodersHugging Face Papers
PAPER·May 28, 2026Self-Improving Language Models with Bidirectional Evolutionary SearchHugging Face Papers
TECH·May 27, 2026Former Google and Apple Researchers Launch a Startup to Build AI’s Missing Feedback LoopWired
PAPER·May 25, 2026From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language ModelsHugging Face Papers
PAPER·May 18, 2026Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy DistillationHugging Face Papers
PAPER·May 13, 2026The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and FixesHugging Face Papers