Switch language한국어
Back to the list

XPENG releases TuringViT for smart driving and humanoid robots

TL;DR AI

Key summary

2 min read
  1. XPENG released TuringViT, a vision encoder for vision-language and vision-language-action systems.

  2. It introduced two versions, TuringViT-18L and TuringViT-24L, aimed at smart driving, cabin AI, and humanoid robots.

  3. XPENG says the smaller model outperformed comparable baselines on throughput at 1536×1536 resolution.

  4. The model was trained on 850 million image-text pairs and reached 83.6% on six zero-shot benchmarks.

Read the original