Improving Model Safety Behavior with Rule-Based Rewards

TL;DR AI
2 min readKey summary
OpenAI outlined Rule-Based Rewards, a training method that uses explicit safety rules instead of relying mainly on human feedback.
The company said the approach has been part of its safety stack since GPT-4 and GPT-4o mini.
It could make safety training faster, easier to update, and more scalable than traditional RLHF.
OpenAI says the method may also help models refuse harmful requests more consistently.



