Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring
TL;DR AI
2 min readKey summary
Researchers benchmarked frontier multimodal LLMs on Clock Drawing Test scoring against supervised vision models.
The models performed reasonably overall, but they showed central tendency bias, pulling predictions toward the middle of the scale.
They mis-scored the most important extreme ratings, which could weaken screening for cognitive impairment.
The study says calibration and post-processing are needed before using LLMs as clinical raters.
