Tsinghua and Alibaba Joint Paper Introduces ViT³: A Vision Transformer with Linear Complexity — CVPR 2026 Oral

TL;DR AI
2 min readKey summary
Tsinghua University and Alibaba researchers introduced ViT³ at CVPR 2026 as a new vision transformer with linear inference complexity.
Instead of standard quadratic attention, ViT³ reframes attention as online test-time training, greatly reducing compute.
The model delivers strong performance across multiple vision tasks while staying more efficient than conventional transformers.
This could make high-resolution vision applications more practical on phones, robots, and other edge devices.



