Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
TL;DR AI
2 min readKey summary
Researchers propose a subject-driven image generation framework that combines multimodal large language models with reference images to jointly encode text and visual cues.
The method adds VAE-based identity conditioning and a Dual Layer Aggregation module with multi-stage denoising.
This helps diffusion models follow text instructions while preserving subject identity more faithfully.
The approach reduces copy-paste artifacts and improves overall image quality.
