PAPER·21 hours agoUltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language ModelsHugging Face Papers
TECH·2 days agoTohoku University and SoftBank move to deploy disaster-prevention AI in societyImpress Watch
PAPER·2 days agoGOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language ModelsarXiv
PAPER·July 17, 2026Show Me Examples: Inferring Visual Concepts from Image SetsApple Machine Learning Research
PAPER·June 2, 2026SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation ModelsHugging Face Papers
PAPER·June 2, 2026PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language UnderstandingHugging Face Papers
PAPER·June 2, 20263DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via CodeHugging Face Papers
PAPER·June 2, 2026VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time OptimizationHugging Face Papers