AI models confidently describe images they never saw, and benchmarks fail to catch it

Key summary
Phantom-0 contains 200 visual questions across 20 categories presented without any image to test image‑less responses.
Several large models (GPT-5, GPT-5.1, GPT-5.2, Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5) confidently described visual details in over 60% of Phantom-0 cases without images.
With typical evaluation prompts that rate rose to 90–100%; overall models achieved 70–80% of benchmark scores without seeing an image, while the actual image added only 20–30%.
Gemini 3 Pro produced diagnoses for nonexistent medical images in five domains (X‑ray, brain MRI, ECG, pathology, dermatology); each medical question was repeated with 200 random seeds. Mirage diagnoses skewed toward severe pathologies (frequent: STEMI, melanomas, carcinomas) though Normal/No diagnosis also appeared.
A failed image upload or agentic/API workflows could lead models to make urgent, incorrect condition recommendations for conditions that do not exist.



