Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer
Исследование пределов трансферного обучения с помощью унифицированного преобразователя текста в текст
2019-10-23
SCID: 54.1/y4yj9rjq
Discuss with AI
Colossal Clean Crawled Corpuslanguage understanding tasksnatural language processingtext-to-text frameworktransfer learning
Figures from the paper
Abstract (AI)
Transfer learning, where a model is first pre-trained on a data-rich task\nbefore being fine-tuned on a downstream task, has emerged as a powerful\ntechnique in natural language processing (NLP). The effectiveness of transfer\nlearning has given rise to a diversity of approaches, methodology, and\npractice. In this paper, we explore the landscape of transfer learning\ntechniques for NLP by introducing a unified framework that converts all\ntext-based language problems into a text-to-text format. Our systematic study\ncompares pre-training objectives, architectures, unlabeled data sets, transfer\napproaches, and other factors on dozens of language understanding tasks. By\ncombining the insights from our exploration with scale and our new ``Colossal\nClean Crawled Corpus'', we achieve state-of-the-art results on many benchmarks\ncovering summarization, question answering, text classification, and more. To\nfacilitate future work on transfer learning for NLP, we release our data set,\npre-trained models, and code.\n
Key Findings
1
A systematic study compares pre-training objectives, model architectures, unlabeled datasets, transfer strategies, and related factors across dozens of language-understanding tasks.
2
Combining the study’s findings with increased model scale and the Colossal Clean Crawled Corpus achieves state-of-the-art results on benchmarks in summarization, question answering, text classification, and other areas.
3
The authors release the Colossal Clean Crawled Corpus, pretrained models, and code to support future NLP transfer-learning research.
4
The paper introduces a unified text-to-text framework that represents diverse NLP problems using a single task format.
Research Object
Unified text-to-text transformer framework for natural language processing (including the Colossal Clean Crawled Corpus and pre-trained models)
Research Subject
the effectiveness and limits of pre-training objectives, architectures, unlabeled datasets, and transfer approaches across diverse NLP tasks
Publication Details
Publication Date
2019-10-23
Journal
Publisher
ISSN
Cited by
8346
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by20
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Transformers: State-of-the-Art Natural Language Processing2020
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing2022
HuggingFace's Transformers: State-of-the-art Natural Language Processing2019
Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Meta LLM-Integrated Systems2020
CodeBERT: A Pre-Trained Model for Programming and Natural Languages2020
ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope2023
Swin Transformer V2: Scaling Up Capacity and Resolution2022
Pre-Trained Image Processing Transformer2021
Video Swin Transformer2022
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing2022
TinyBERT: Distilling BERT for Natural Language Understanding2020
Lost in the Middle: How Language Models Use Long Contexts2024
Efficient Transformers: A Survey2022
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021
A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly2024
GLM: General Language Model Pretraining with Autoregressive Blank Infilling2022
Large Language Models for Software Engineering: A Systematic Literature Review2024
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models2024
A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges2022