Switch language한국어
Back to the list

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

TL;DR AI

Key summary

2 min read
  1. Lens is a compact 3.8B text-to-image model that achieves competitive or better results than larger systems while using only about 19.3% of Z-Image’s training compute.

  2. It was trained on the Lens-800M dataset with dense captions, multi-resolution batching, and other efficiency-focused optimization techniques.

  3. The model supports multilingual generation and multiple aspect ratios, making it more flexible for real-world use.

  4. Lens also includes RL fine-tuning and a 4-step distilled turbo version for faster inference.

Read the original