PAPER·14 hours agoCLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge AcquisitionHugging Face Papers
PAPER·2 days agoMage-VL: An Efficient Codec-Native Streaming Multimodal Foundation ModelHugging Face Papers
TECH·July 24, 2026Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to startVentureBeat
PAPER·July 23, 2026MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement JudgementarXiv
TECH·July 20, 2026Bilibili showcases N.E.K.O., an AI companion that can interpret desktop content and initiate conversationsTechNode