Pinterest cut AI costs 90% on Qwen fine-tuning

TL;DR AI
2 min readKey summary
Pinterest says it cut AI costs by 90% and boosted accuracy by 30% by customizing Qwen3-VL for shopping and discovery.
The company removed the model’s vision encoder and replaced it with proprietary multimodal embeddings, enabling offline image metadata precomputation and faster runtime inference.
Pinterest can now retrain more often, reduce expensive per-image calls, and improve latency across visual search and recommendations.
It also uses a taste graph and user embeddings to track evolving preferences and strengthen personalized discovery at scale.
