Switch language한국어
Back to the list

Frontier AI Models Are Doing Something Absolutely Bizarre When Asked to Diagnose Medical X-Rays

TL;DR AI

Key summary

2 min read
  1. Stanford researchers found that leading multimodal AI models can sound accurate on image questions even when they are not shown the image.

  2. The team called this behavior “mirage reasoning,” noting that models like GPT-5, Gemini 3 Pro, and Claude Opus 4.5 can generate convincing but false visual descriptions.

  3. On a chest X-ray task and other benchmarks, the effect can inflate scores by rewarding pattern-based guesswork instead of real vision.

  4. The finding raises safety concerns for medical imaging, where overestimated performance could hide serious diagnostic risks.

Read the original