Switch language한국어
Back to the list

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced χ-Bench, a benchmark for healthcare workflows spanning provider prior authorization, payer utilization management, and care management.

  2. The best-performing agent completed only 28.0% of tasks across model and harness setups.

  3. Performance dropped sharply when tasks had to be completed in a single session, highlighting weak long-horizon execution.

  4. The results point to major gaps in AI agents handling policy-heavy, multi-step enterprise workflows, especially in regulated settings.

Read the original