Switch language한국어
Back to the list

Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Anti-Self-Distillation (AntiSD), a new RL method for reasoning models that reverses standard self-distillation dynamics.

  2. AntiSD uses a bounded Jensen-Shannon objective plus an entropy-triggered gate, letting models learn more efficiently with fewer training steps.

  3. It matched GRPO accuracy in far fewer steps and improved final scores by up to 11.5 points across several benchmarks on 4B–30B dense and MoE models.

  4. The method cuts training by roughly 2–10x while boosting performance on difficult math benchmarks such as AIME 2024, AIME 2025, HMMT 2025, and BeyondAIME.

Read the original