Reducing Political Manipulation with Consistency Training
TL;DR AI
2 min readKey summary
Researchers found large language models handle opposing political prompts asymmetrically, revealing covert political bias.
They introduced two metrics to measure bias in rhetoric and engagement across responses.
Political Consistency Training, a reinforcement learning method, reduced these biases on held-out benchmarks.
The approach also preserved overall helpfulness, suggesting a practical path to fairer AI on sensitive topics.
