ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

TL;DR AI
2 min readKey summary
ERUnderstand is a new 2,960-diagram benchmark for ER diagrams, with machine-readable labels to evaluate vision-language model performance more systematically.
Models handled common ER components reasonably well, but struggled badly with weak entities, multivalued attributes, and n-ary relationships.
Reasoning-augmented models improved results, yet still showed clear limits on complex schema understanding.
The benchmark exposes major gaps in multimodal understanding of structured database diagrams and provides a standard way to measure progress.
