RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

TL;DR AI
2 min readKey summary
Researchers introduced RealICU, a hindsight-labeled benchmark for testing LLM agents on long ICU patient trajectories and clinical decision support.
RealICU uses full-context review by senior physicians and includes tasks for patient status, acute problems, recommended actions, and unsafe red-flag actions.
It ships as RealICU-Gold and a larger RealICU-Scale, built from ICU trajectories derived from MIMIC-IV and ICU-Evo.
Tests showed current LLMs performed poorly, often anchoring to early impressions and facing a recall-versus-safety tradeoff.
A structured-memory agent improved long-range reasoning, but safety problems still remained.
