Same prompt, different morals: how frontier AI models diverge on ethical dilemmas

TL;DR AI
2 min readKey summary
Philosophy Bench tested leading AI models on 100 everyday ethical dilemmas and found clear differences in moral style.
Claude was the most deontological and refusal-prone, while Grok was the most consequentialist and compliant.
Gemini was easiest to steer, and GPT-5 family models made fewer mistakes but used little explicit moral language.
The results suggest frontier models are already shipping with distinct ethical behaviors that could matter more as agentic AI grows more powerful.
