Switch language한국어
Back to the list

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced UltraViT, a vision encoder designed specifically to reduce on-device latency for large vision-language models.

  2. It uses a pyramidal architecture with heterogeneous spatial mixers, plus a two-stage generative pre-training scheme.

  3. The pre-training combines dense distillation and frozen-LLM supervision to improve multimodal alignment and efficiency.

  4. In experiments, UltraViT beat encoder-focused baselines while running about 1.7× faster on device.

Read the original