Show Me Examples: Inferring Visual Concepts from Image Sets

TL;DR AI
2 min readKey summary
Researchers introduce VICIS, a new task for visual concept inference from image sets in vision-language models.
The paper finds current VLMs struggle to identify a shared concept from examples and apply it to a query image.
They propose a model and training framework that learn concept-specific embeddings for better generation.
The approach improves quality, diversity, and generalization on synthetic, ImageNet/WordNet, and sketch data.
