Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving
TL;DR AI
2 min readKey summary
Researchers introduced Fast-dDrive, a block-diffusion vision-language-action model for autonomous driving.
It enforces causal structure across semantic output blocks, freezes structural tokens, and uses scaffold speculative decoding plus shared-prefix rollouts to cut error and boost efficiency.
The model reports strong results on WOD-E2E and nuScenes, and delivers a large throughput gain over an autoregressive baseline when paired with SGLang.
The approach improves real-time deployment viability by speeding up inference without sacrificing trajectory quality.
