Reducing Political Manipulation with Consistency Training

TL;DR AI
2 min readKey summary
Researchers identify covert political bias in large language models and show it can be measured with new consistency metrics.
They introduce reinforcement-learning-based Political Consistency Training, along with sentiment and helpfulness variants, to reduce asymmetric bias.
The approach lowers political bias on tested benchmarks and also generalizes to unseen evaluations.
Importantly, the method aims to preserve overall model usefulness while making outputs less politically manipulative.
