Switch language한국어
Back to the list

AI models confidently describe images they never saw, and benchmarks fail to catch it

TL;DR AI

Key summary

2 min read
  1. Phantom-0 contains 200 visual questions across 20 categories presented without any image to test image‑less responses.

  2. Several large models (GPT-5, GPT-5.1, GPT-5.2, Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5) confidently described visual details in over 60% of Phantom-0 cases without images.

  3. With typical evaluation prompts that rate rose to 90–100%; overall models achieved 70–80% of benchmark scores without seeing an image, while the actual image added only 20–30%.

  4. Gemini 3 Pro produced diagnoses for nonexistent medical images in five domains (X‑ray, brain MRI, ECG, pathology, dermatology); each medical question was repeated with 200 random seeds. Mirage diagnoses skewed toward severe pathologies (frequent: STEMI, melanomas, carcinomas) though Normal/No diagnosis also appeared.

  5. A failed image upload or agentic/API workflows could lead models to make urgent, incorrect condition recommendations for conditions that do not exist.

Read the original