Switch language한국어
Back to the list

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced SpatialUncertain, a benchmark for testing uncertainty handling in vision-language models on spatial reasoning tasks.

  2. Across frontier open- and closed-source models, accuracy dropped sharply under occlusion and perspective ambiguity.

  3. Models often failed to abstain when evidence was insufficient and were poor at identifying useful additional viewpoints.

  4. The findings show vision-language models can be confidently wrong on spatial questions, raising safety concerns for real-world use.

Read the original