PAPER·3 days agoX-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted MatchingarXiv
PAPER·3 days agoEchoverse: Deep, Evolving Environments for Training Computer-Use Agents at ScalearXiv
PAPER·5 days agoCollaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement LearningarXiv
PAPER·6 days agoSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving SkillsHugging Face Papers
PAPER·6 days agoMolt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningHugging Face Papers