OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation
TL;DR AI
2 min readKey summary
Researchers introduced OcclusionFormer, a diffusion-transformer method for layout-to-image generation that explicitly models object depth order and occlusion.
The paper also presents SA-Z, a dataset with pixel-level annotations and explicit Z-order labels for overlapping instances.
OcclusionFormer separates objects and composites them with volume rendering, guided by an alignment loss to improve layered scene synthesis.
The approach addresses a major failure mode in controllable image generation: handling overlapping objects in a physically plausible way.
