Switch language한국어
Back to the list

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

TL;DR AI

Key summary

2 min read
  1. Researchers introduced SafeSteer, a localized on-policy distillation method for LLM safety alignment.

  2. It builds a safety teacher with activation steering, identifies safety tokens, and applies reverse KL only to those tokens.

  3. In tests, SafeSteer delivered strong safety gains with only small capability drops, using just 100 harmful samples and no general-purpose data.

  4. The approach could reduce the data and training cost of alignment while limiting the alignment tax on general performance.

Read the original