Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
TL;DR AI
2 min readKey summary
Lens is a compact 3.8B text-to-image model that achieves competitive or better results than larger systems while using only about 19.3% of Z-Image’s training compute.
It was trained on the Lens-800M dataset with dense captions, multi-resolution batching, and other efficiency-focused optimization techniques.
The model supports multilingual generation and multiple aspect ratios, making it more flexible for real-world use.
Lens also includes RL fine-tuning and a 4-step distilled turbo version for faster inference.
