Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
TL;DR AI
2 min readKey summary
Researchers introduced PALATE, a benchmark for evaluating role-playing chatbots.
PALATE assesses agents through conversations with simulated users built from 300 character profiles.
It goes beyond fixed dialogue histories and generic rubrics by supporting personalized evaluation.
The benchmark was used to test 16 candidate systems and measure turn quality, long-horizon performance, and user-specific satisfaction.
