GenTron: Diffusion Transformers for Image and Video Generation

GenTron: диффузионные трансформеры для генерации изображений и видео
Ping Luo, Tao Xiang, Mengmeng Xu, Shoufa Chen, Jiawei Ren, Yuren Cong, Sen He, Yanping Xie, Animesh A. Sinha, Juan-Manuel Pérez-Rúa
2024-06-16

Diffusion TransformersGenTronTransformer-based diffusionmotion-free guidancetext-to-video generation
In this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexibility and scalability, the visual generative domain primarily utilizes CNN-based U-Net architectures, particularly in diffusion-based models. We introduce GenTron, a family of Generative models employing Transformer-based diffusion, to address this gap. Our initial step was to adapt Diffusion Transformers (DiTs) from class to text conditioning, a process involving thorough empirical exploration of the conditioning mechanism. We then scale GenTron from approximately 900M to over 3B parameters, observing improvements in visual quality. Furthermore, we extend GenTron to text-to-video generation, incorporating novel motion-free guidance to enhance video quality. In human evaluations against SDXL, GenTron achieves a 51.1% win rate in visual quality (with a 19.8% draw rate), and a 42.3% win rate in text alignment (with a 42.9% draw rate). GenTron notably performs well in T2I-CompBench, highlighting its compositional generation ability. We hope GenTron could provide meaningful insights and serve as a valuable reference for future research. Please refer to the website<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://www.shoufachen.com/gentron_website/ and the arXiv version for the most up-to-date results: https://arxiv.org/abs/2312.04557.
1
Adaptation of Diffusion Transformers (DiTs) from class conditioning to text conditioning was performed through extensive empirical exploration of conditioning mechanisms.
2
GenTron introduces a family of Transformer-based diffusion generative models for image and video, addressing reliance on CNN U-Net architectures.
3
GenTron is extended to text-to-video generation and employs a novel motion-free guidance technique to improve video quality.
4
GenTron performs notably well on T2I-CompBench, demonstrating strong compositional text-to-image generation ability.
5
Human evaluations against SDXL show GenTron achieves a 51.1% win rate in visual quality (19.8% draws) and a 42.3% win rate in text alignment (42.9% draws).
6
Scaling GenTron from ~900M to over 3B parameters improves visual quality, indicating positive scaling behavior.

GenTron family of Transformer-based diffusion generative models for image and video generation

Use and evaluation of Transformer-based diffusion architectures (scaling from ~900M to >3B parameters), conditioning mechanisms for text conditioning, and motion-free guidance for text-to-video to improve visual quality and text alignment

Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Cited by
38
Access Type
Author Information
Authors
Ping Luo
Tao Xiang
Mengmeng Xu
Shoufa Chen
Jiawei Ren
Yuren Cong
Sen He
Yanping Xie
Animesh A. Sinha
Juan-Manuel Pérez-Rúa
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%