Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
TL;DR AI
2 min readKey summary
Researchers introduced Grounded Personality Reasoning and the MM-OCEAN benchmark to test whether multimodal models can justify Big Five personality ratings from observable cues.
Across 27 models, many correct personality predictions were not actually grounded in evidence, revealing high prejudice and confabulation rates.
The results highlight a gap between accurate judgments and evidence-based reasoning, raising safety and reliability concerns for human-facing AI.
