Taming Pretrained Transformers for Extreme Multi-label Text Classification

Адаптация предобученных трансформеров для экстремальной многометочной классификации текстов
Wei-Cheng Chang, Hsiang‐Fu Yu, Kai Zhong, Yiming Yang, Inderjit S. Dhillon
2020-08-20

Amazon product2queryX-Transformerextreme multi-label text classificationlabel sparsitypretrained transformers
We consider the extreme multi-label text classification (XMC) problem: given an input text, return the most relevant labels from a large label collection. For example, the input text could be a product description on Amazon.com and the labels could be product categories. XMC is an important yet challenging problem in the NLP community. Recently, deep pretrained transformer models have achieved state-of-the-art performance on many NLP tasks including sentence classification, albeit with small label sets. However, naively applying deep transformer models to the XMC problem leads to sub-optimal performance due to the large output space and the label sparsity issue. In this paper, we propose X-Transformer, the first scalable approach to fine-tuning deep transformer models for the XMC problem. The proposed method achieves new state-of-the-art results on four XMC benchmark datasets. In particular, on a Wiki dataset with around 0.5 million labels, the [email protected] of X-Transformer is 77.28%, a substantial improvement over state-of-the-art XMC approaches Parabel (linear) and AttentionXML (neural), which achieve 68.70% and 76.95% [email protected], respectively. We further apply X-Transformer to a product2query dataset from Amazon and gained 10.7% relative improvement on [email protected] over Parabel.
1
Naively applying pretrained transformers to XMC is suboptimal because of the large output space and sparse labels.
2
On a Wiki dataset with approximately 0.5 million labels, X-Transformer reaches 77.28% P@1, versus 68.70% for Parabel and 76.95% for AttentionXML.
3
On an Amazon product2query dataset, X-Transformer improves P@1 by 10.7% relative to Parabel.
4
The paper introduces X-Transformer, the first scalable method for fine-tuning deep pretrained transformers for extreme multi-label text classification.
5
X-Transformer achieves state-of-the-art results on four XMC benchmark datasets.

Extreme multi-label text classification of input texts against large label collections

Scalable fine-tuning of pretrained transformer models to address large output spaces and label sparsity, evaluated by XMC ranking performance

Publication Details
Publication Date
2020-08-20
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Wei-Cheng Chang
Hsiang‐Fu Yu
Kai Zhong
Yiming Yang
Inderjit S. Dhillon
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%