CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
TL;DR AI
2 min readKey summary
Researchers introduced χ-Bench, a benchmark for healthcare workflows spanning provider prior authorization, payer utilization management, and care management.
The best-performing agent completed only 28.0% of tasks across model and harness setups.
Performance dropped sharply when tasks had to be completed in a single session, highlighting weak long-horizon execution.
The results point to major gaps in AI agents handling policy-heavy, multi-step enterprise workflows, especially in regulated settings.
