Switch language한국어
Back to the list

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced RealICU, a hindsight-labeled benchmark for testing LLM agents on long ICU patient trajectories and clinical decision support.

  2. RealICU uses full-context review by senior physicians and includes tasks for patient status, acute problems, recommended actions, and unsafe red-flag actions.

  3. It ships as RealICU-Gold and a larger RealICU-Scale, built from ICU trajectories derived from MIMIC-IV and ICU-Evo.

  4. Tests showed current LLMs performed poorly, often anchoring to early impressions and facing a recall-versus-safety tradeoff.

  5. A structured-memory agent improved long-range reasoning, but safety problems still remained.

Read the original