Vanilla ViT for Automotive Point Cloud Semantic Segmentation

TL;DR AI
2 min readKey summary
Researchers introduced VaViT, a vanilla Vision Transformer for automotive LiDAR point cloud semantic segmentation.
With a redesigned tokenizer, lightweight decoder, and targeted augmentations, it closes much of the gap with U-Net-style models.
VaViT delivers strong results on nuScenes, SemanticKITTI, and Waymo Open Dataset.
The work suggests plain ViTs can compete with specialized 3D segmentation architectures, simplifying autonomous driving perception stacks.
