PAPER·June 1, 2026The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised ImprovementHugging Face Papers
PAPER·May 29, 2026Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned BiasesHugging Face Papers
PAPER·May 27, 2026Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned BiasesarXiv
TECH·May 25, 2026StepFun Releases StepAudio 2.5 Realtime: An End-to-End Voice Model with Roleplay-Specific RLHF and Paralinguistic ComprehensionMarkTechPost
PAPER·May 21, 2026Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable AlignmentHugging Face Papers
CODING·May 17, 2026Understanding Reinforcement Learning with Neural Networks Part 6: Completing the Reinforcement Learning ProcessDev.to
TECH·May 1, 2026Why OpenAI's 'goblin' problem matters — and how you can release the goblins on your ownVentureBeat