Making AI chatbots helpful weakens their ability to simulate human behavior, large-scale study finds

TL;DR AI
2 min readKey summary
A large study found that base language models matched human responses better than their post-trained chatbot versions.
The biggest drop came after reasoning training, followed by instruction tuning and vision-related post-training.
The gap grew across newer generations of Qwen, Llama, and OLMo models.
Adding participant-specific role information did not fix the mismatch.


