Switch language한국어
Back to the list

OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OcclusionFormer, a diffusion-transformer method for layout-to-image generation that explicitly models object depth order and occlusion.

  2. The paper also presents SA-Z, a dataset with pixel-level annotations and explicit Z-order labels for overlapping instances.

  3. OcclusionFormer separates objects and composites them with volume rendering, guided by an alignment loss to improve layered scene synthesis.

  4. The approach addresses a major failure mode in controllable image generation: handling overlapping objects in a physically plausible way.

Read the original