QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
TL;DR AI
2 min readKey summary
Researchers introduced QUACK, an open-source benchmark and auditing framework for multimodal social deduction agents.
QUACK evaluates agents at three levels: final outcomes, behavioral trajectories, and individual utterances.
Tests on frontier vision-language models found frequent hallucinations and accusations made without supporting evidence.
The work argues that agent evaluation should measure grounding and consistency, not just whether a model wins.
