HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips

HowTo100M: обучение текстово-видеового представления посредством просмотра ста миллионов озвученных видеоклипов
Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, Josef Šivic, Dimitri Zhukov, Makarand Tapaswi
2019-06-07

HowTo100M datasetaction localizationautomatically transcribed narrationstext-to-video retrievaltext-video embedding
Learning text-video embeddings usually requires a dataset of video clips with\nmanually provided captions. However, such datasets are expensive and time\nconsuming to create and therefore difficult to obtain on a large scale. In this\nwork, we propose instead to learn such embeddings from video data with readily\navailable natural language annotations in the form of automatically transcribed\nnarrations. The contributions of this work are three-fold. First, we introduce\nHowTo100M: a large-scale dataset of 136 million video clips sourced from 1.22M\nnarrated instructional web videos depicting humans performing and describing\nover 23k different visual tasks. Our data collection procedure is fast,\nscalable and does not require any additional manual annotation. Second, we\ndemonstrate that a text-video embedding trained on this data leads to\nstate-of-the-art results for text-to-video retrieval and action localization on\ninstructional video datasets such as YouCook2 or CrossTask. Finally, we show\nthat this embedding transfers well to other domains: fine-tuning on generic\nYoutube videos (MSR-VTT dataset) and movies (LSMDC dataset) outperforms models\ntrained on these datasets alone. Our dataset, code and models will be publicly\navailable at: www.di.ens.fr/willow/research/howto100m/.\n
1
A text-video embedding trained on HowTo100M achieves state-of-the-art text-to-video retrieval and action localization on YouCook2 and CrossTask.
2
HowTo100M introduces 136 million video clips from 1.22 million narrated instructional web videos covering over 23,000 visual tasks.
3
The dataset uses automatically transcribed narrations as natural-language annotations, enabling fast, scalable collection without additional manual labeling.
4
The learned embedding transfers effectively across domains, outperforming models trained only on MSR-VTT or LSMDC after fine-tuning.
5
The work provides a publicly available dataset, code, and trained models for large-scale text-video representation learning.

HowTo100M: 136 million narrated video clips from 1.22 million instructional web videos depicting humans performing and describing over 23,000 visual tasks

Learning a text-video embedding from automatically transcribed narrations and its transfer performance for text-to-video retrieval and action localization across video domains

Publication Details
Publication Date
2019-06-07
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Jean-Baptiste Alayrac
Antoine Miech
Ivan Laptev
Josef Šivic
Dimitri Zhukov
Makarand Tapaswi
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%