GenTron: Diffusion Transformers for Image and Video Generation
GenTron: диффузионные трансформеры для генерации изображений и видео
2024-06-16
SCID: 54.1/syp3huy5
Discuss with AI
Diffusion TransformersGenTronTransformer-based diffusionmotion-free guidancetext-to-video generation
Figures from the paper
Abstract (AI)
In this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexibility and scalability, the visual generative domain primarily utilizes CNN-based U-Net architectures, particularly in diffusion-based models. We introduce GenTron, a family of Generative models employing Transformer-based diffusion, to address this gap. Our initial step was to adapt Diffusion Transformers (DiTs) from class to text conditioning, a process involving thorough empirical exploration of the conditioning mechanism. We then scale GenTron from approximately 900M to over 3B parameters, observing improvements in visual quality. Furthermore, we extend GenTron to text-to-video generation, incorporating novel motion-free guidance to enhance video quality. In human evaluations against SDXL, GenTron achieves a 51.1% win rate in visual quality (with a 19.8% draw rate), and a 42.3% win rate in text alignment (with a 42.9% draw rate). GenTron notably performs well in T2I-CompBench, highlighting its compositional generation ability. We hope GenTron could provide meaningful insights and serve as a valuable reference for future research. Please refer to the website<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://www.shoufachen.com/gentron_website/ and the arXiv version for the most up-to-date results: https://arxiv.org/abs/2312.04557.
Key Findings
1
Adaptation of Diffusion Transformers (DiTs) from class conditioning to text conditioning was performed through extensive empirical exploration of conditioning mechanisms.
2
GenTron introduces a family of Transformer-based diffusion generative models for image and video, addressing reliance on CNN U-Net architectures.
3
GenTron is extended to text-to-video generation and employs a novel motion-free guidance technique to improve video quality.
4
GenTron performs notably well on T2I-CompBench, demonstrating strong compositional text-to-image generation ability.
5
Human evaluations against SDXL show GenTron achieves a 51.1% win rate in visual quality (19.8% draws) and a 42.3% win rate in text alignment (42.9% draws).
6
Scaling GenTron from ~900M to over 3B parameters improves visual quality, indicating positive scaling behavior.
Research Object
GenTron family of Transformer-based diffusion generative models for image and video generation
Research Subject
Use and evaluation of Transformer-based diffusion architectures (scaling from ~900M to >3B parameters), conditioning mechanisms for text conditioning, and motion-free guidance for text-to-video to improve visual quality and text alignment
Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Cited by
38
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai8
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
LLaMA: Open and Efficient Foundation Language Models2023
PaLM: Scaling Language Modeling with Pathways2022
Scalable Diffusion Models with Transformers2023
Scaling Vision Transformers2022
ImageNet: A large-scale hierarchical image database2009