Switch language한국어
Back to the list

Making AI chatbots helpful weakens their ability to simulate human behavior, large-scale study finds

TL;DR AI

Key summary

2 min read
  1. A large study found that base language models matched human responses better than their post-trained chatbot versions.

  2. The biggest drop came after reasoning training, followed by instruction tuning and vision-related post-training.

  3. The gap grew across newer generations of Qwen, Llama, and OLMo models.

  4. Adding participant-specific role information did not fix the mismatch.

Read the original