The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

TL;DR AI
2 min readKey summary
OpenAI says it trained GPT-3.5 to follow an instruction hierarchy that separates privileged commands from untrusted prompts.
The approach is designed to help the model ignore malicious prompt injections and jailbreak attempts.
Using generated training data, OpenAI says the model became more robust across both seen and unseen adversarial prompts.
The company says the security gains come without a major drop in normal performance.



