Switch language한국어
Back to the list

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OGPSA, an orthogonal gradient projection method for safety alignment in large language models.

  2. It treats sequential safety training as continual learning and projects safety gradients away from a reference subspace.

  3. Across SFT, DPO, and combined pipelines, OGPSA improves policy compliance while better preserving general capabilities.

  4. The approach aims to reduce alignment tax and gradient interference, a key challenge in LLM post-training.

Read the original