Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
TL;DR AI
2 min readKey summary
Researchers introduced OGPSA, an orthogonal gradient projection method for safety alignment in large language models.
It treats sequential safety training as continual learning and projects safety gradients away from a reference subspace.
Across SFT, DPO, and combined pipelines, OGPSA improves policy compliance while better preserving general capabilities.
The approach aims to reduce alignment tax and gradient interference, a key challenge in LLM post-training.
