Switch language한국어
Back to the list

Improving Model Safety Behavior with Rule-Based Rewards

TL;DR AI

Key summary

2 min read
  1. OpenAI outlined Rule-Based Rewards, a training method that uses explicit safety rules instead of relying mainly on human feedback.

  2. The company said the approach has been part of its safety stack since GPT-4 and GPT-4o mini.

  3. It could make safety training faster, easier to update, and more scalable than traditional RLHF.

  4. OpenAI says the method may also help models refuse harmful requests more consistently.

Read the original