How can embedding models bind concepts?
TL;DR AI
2 min readKey summary
Researchers studied concept binding in vision-language embedding models and how objects are combined into scenes.
CLIP can recover object information from separate embeddings, but it still struggles with binding concepts compositionally.
The paper argues CLIP’s binding function is high-complexity, which limits shared generalization across modalities.
By contrast, transformers trained from scratch learned simpler multiplicative binding functions and generalized better with enough data.
