BloombergGPT: A Large Language Model for Finance

BloombergGPT: большая языковая модель для финансов
Ozan İrsoy, Sebastian Gehrmann, Mark Dredze, Steven Lu, Prabhanjan Kambadur, Shijie Wu, Vadim Dabravolski, David Rosenberg, Gideon Mann
2023-03-30

50 billion parameter language modelBloombergGPTfinancial NLP benchmarksfinancial domain dataset (363B tokens)mixed dataset training (financial + general 345B tokens)
The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering. Large Language Models (LLMs) have been shown to be effective on a variety of tasks; however, no LLM specialized for the financial domain has been reported in literature. In this work, we present BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data. We construct a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet, augmented with 345 billion tokens from general purpose datasets. We validate BloombergGPT on standard LLM benchmarks, open financial benchmarks, and a suite of internal benchmarks that most accurately reflect our intended usage. Our mixed dataset training leads to a model that outperforms existing models on financial tasks by significant margins without sacrificing performance on general LLM benchmarks. Additionally, we explain our modeling choices, training process, and evaluation methodology. We release Training Chronicles (Appendix C) detailing our experience in training BloombergGPT.
1
BloombergGPT is a 50 billion parameter large language model specifically trained for the financial domain.
2
BloombergGPT maintains competitive performance on standard general LLM benchmarks while improving financial task performance.
3
Mixed-domain training yields a model that outperforms existing models on financial tasks by significant margins.
4
The paper documents modeling choices, training process, evaluation methodology, and provides detailed Training Chronicles (Appendix C).
5
Training corpus comprises a 363 billion token finance-specific dataset plus 345 billion tokens from general-purpose datasets.

BloombergGPT, a 50-billion-parameter large language model trained on financial and general-purpose text

Performance and behavior of BloombergGPT on financial NLP tasks and general LLM benchmarks as influenced by mixed-domain (financial + general) large-scale training data and training choices

Publication Details
Publication Date
2023-03-30
Journal
Publisher
ISSN
Cited by
309
Access Type
Author Information
Authors
Ozan İrsoy
Sebastian Gehrmann
Mark Dredze
Steven Lu
Prabhanjan Kambadur
Shijie Wu
Vadim Dabravolski
David Rosenberg
Gideon Mann
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%