SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

TL;DR AI
2 min readKey summary
Researchers introduced SafeSteer, a localized on-policy distillation method for LLM safety alignment.
It builds a safety teacher with activation steering, identifies safety tokens, and applies reverse KL only to those tokens.
In tests, SafeSteer delivered strong safety gains with only small capability drops, using just 100 harmful samples and no general-purpose data.
The approach could reduce the data and training cost of alignment while limiting the alignment tax on general performance.
