Switch language한국어
Back to the list

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

TL;DR AI

Key summary

2 min read
  1. OpenAI says it trained GPT-3.5 to follow an instruction hierarchy that separates privileged commands from untrusted prompts.

  2. The approach is designed to help the model ignore malicious prompt injections and jailbreak attempts.

  3. Using generated training data, OpenAI says the model became more robust across both seen and unseen adversarial prompts.

  4. The company says the security gains come without a major drop in normal performance.

Read the original