Naver's Seoul World Model uses Street View images to prevent AI from hallucinating cities

TL;DR AI
2 min readKey summary
Naver built the Seoul World Model using over a million Street View panoramas to ground video generation in real city geometry.
The model trains with cross-temporal pairing and synthetic CARLA videos to avoid copying transient objects and to fill missing viewpoints.
A virtual lookahead sink provides moving landmarks to prevent drift over long routes while depth and latent encodings supply layout and appearance.
SWM was trained on 24 Nvidia H100 GPUs and generalizes to unseen cities, outperforming prior world models on multiple benchmarks.



