HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
HowTo100M: обучение текстово-видеового представления посредством просмотра ста миллионов озвученных видеоклипов
2019-06-07
SCID: 54.1/mcpareua
Discuss with AI
HowTo100M datasetaction localizationautomatically transcribed narrationstext-to-video retrievaltext-video embedding
Figures from the paper
Abstract (AI)
Learning text-video embeddings usually requires a dataset of video clips with\nmanually provided captions. However, such datasets are expensive and time\nconsuming to create and therefore difficult to obtain on a large scale. In this\nwork, we propose instead to learn such embeddings from video data with readily\navailable natural language annotations in the form of automatically transcribed\nnarrations. The contributions of this work are three-fold. First, we introduce\nHowTo100M: a large-scale dataset of 136 million video clips sourced from 1.22M\nnarrated instructional web videos depicting humans performing and describing\nover 23k different visual tasks. Our data collection procedure is fast,\nscalable and does not require any additional manual annotation. Second, we\ndemonstrate that a text-video embedding trained on this data leads to\nstate-of-the-art results for text-to-video retrieval and action localization on\ninstructional video datasets such as YouCook2 or CrossTask. Finally, we show\nthat this embedding transfers well to other domains: fine-tuning on generic\nYoutube videos (MSR-VTT dataset) and movies (LSMDC dataset) outperforms models\ntrained on these datasets alone. Our dataset, code and models will be publicly\navailable at: www.di.ens.fr/willow/research/howto100m/.\n
Key Findings
1
A text-video embedding trained on HowTo100M achieves state-of-the-art text-to-video retrieval and action localization on YouCook2 and CrossTask.
2
HowTo100M introduces 136 million video clips from 1.22 million narrated instructional web videos covering over 23,000 visual tasks.
3
The dataset uses automatically transcribed narrations as natural-language annotations, enabling fast, scalable collection without additional manual labeling.
4
The learned embedding transfers effectively across domains, outperforming models trained only on MSR-VTT or LSMDC after fine-tuning.
5
The work provides a publicly available dataset, code, and trained models for large-scale text-video representation learning.
Research Object
HowTo100M: 136 million narrated video clips from 1.22 million instructional web videos depicting humans performing and describing over 23,000 visual tasks
Research Subject
Learning a text-video embedding from automatically transcribed narrations and its transfer performance for text-to-video retrieval and action localization across video domains
Publication Details
Publication Date
2019-06-07
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest