Switch language한국어
Back to the list

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring

TL;DR AI

Key summary

2 min read
  1. Researchers benchmarked frontier multimodal LLMs on Clock Drawing Test scoring against supervised vision models.

  2. The models performed reasonably overall, but they showed central tendency bias, pulling predictions toward the middle of the scale.

  3. They mis-scored the most important extreme ratings, which could weaken screening for cognitive impairment.

  4. The study says calibration and post-processing are needed before using LLMs as clinical raters.

Read the original