RenAIssance: A Survey Into AI Text-to-Image Generation in the Era of Large Model
RenAIssance: Обзор генерации изображений по тексту на основе ИИ в эпоху больших моделей
2024-12-27
SCID: 54.1/t9zre2kq
Discuss with AI
Generative Adversarial Network (GAN)diffusion modelslarge language models (LLMs)scaling large modelstext-to-image generation
Figures from the paper
Abstract (AI)
Text-to-image generation (TTI) refers to the usage of models that could process text input and generate high fidelity images based on text descriptions. Text-to-image generation using neural networks could be traced back to the emergence of Generative Adversial Network (GAN), followed by the autoregressive Transformer. Diffusion models are one prominent type of generative model used for the generation of images through the systematic introduction of noises with repeating steps. As an effect of the impressive results of diffusion models on image synthesis, it has been cemented as the major image decoder used by text-to-image models and brought text-to-image generation to the forefront of machine-learning (ML) research. In the era of large models, scaling up model size and the integration with large language models have further improved the performance of TTI models, resulting the generation result nearly indistinguishable from real-world images, revolutionizing the way we retrieval images. Our explorative study has incentivised us to think that there are further ways of scaling text-to-image models with the combination of innovative model architectures and prediction enhancement techniques. We have divided the work of this survey into five main sections wherein we detail the frameworks of major literature in order to delve into the different types of text-to-image generation methods. Following this we provide a detailed comparison and critique of these methods and offer possible pathways of improvement for future work. In the future work, we argue that TTI development could yield impressive productivity improvements for creation, particularly in the context of the AIGC era, and could be extended to more complex tasks such as video generation and 3D generation.
Key Findings
1
Authors identify opportunities for further scaling via novel model architectures and prediction enhancement techniques to improve TTI models.
2
Diffusion models have become the dominant image decoder in text-to-image (TTI) generation due to their impressive image synthesis results.
3
Scaling up model size and integrating with large language models significantly improved TTI performance, producing results nearly indistinguishable from real-world images.
4
TTI advances are expected to boost creative productivity in the AIGC era and to be extendable to more complex tasks like video and 3D generation.
5
The survey categorizes TTI literature into five sections detailing major frameworks, methods, and architectures for text-to-image generation.
Research Object
Text-to-image generation models (AI models that convert text descriptions into images)
Research Subject
Techniques, architectures, scaling and integration with large language models, and performance/quality of generated images in modern text-to-image generation
Publication Details
Publication Date
2024-12-27
Journal
Publisher
ISSN
Cited by
54
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai14
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
ImageNet: A large-scale hierarchical image database2009
Learning Multiple Layers of Features from Tiny Images2024
Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks2017
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network2017
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
LLaMA: Open and Efficient Foundation Language Models2023
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers2021
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
PaLM: Scaling Language Modeling with Pathways2022
A Survey of Large Language Models2026
Classifier-Free Diffusion Guidance2022
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2023
BloombergGPT: A Large Language Model for Finance2023