LaCoVL-FER: Landmark-Guided Contrastive Learning Network with Vision-Language Enhancement for Facial Expression Recognition

TL;DR AI
2 min readKey summary
Researchers introduced LaCoVL-FER, a facial expression recognition model that combines facial landmarks with CLIP-based vision-language features.
The method uses landmark-guided adaptive encoding and vision-language enhancement to fuse geometric and semantic cues.
On RAF-DB, FERPlus, and AffectNet, the model delivered state-of-the-art performance.
The approach is designed to improve robustness under pose changes, occlusion, and lighting variation.
