BloombergGPT: A Large Language Model for Finance
BloombergGPT: большая языковая модель для финансов
2023-03-30
SCID: 54.1/8auhbq5n
Discuss with AI
50 billion parameter language modelBloombergGPTfinancial NLP benchmarksfinancial domain dataset (363B tokens)mixed dataset training (financial + general 345B tokens)
Figures from the paper
Abstract (AI)
The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering. Large Language Models (LLMs) have been shown to be effective on a variety of tasks; however, no LLM specialized for the financial domain has been reported in literature. In this work, we present BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data. We construct a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet, augmented with 345 billion tokens from general purpose datasets. We validate BloombergGPT on standard LLM benchmarks, open financial benchmarks, and a suite of internal benchmarks that most accurately reflect our intended usage. Our mixed dataset training leads to a model that outperforms existing models on financial tasks by significant margins without sacrificing performance on general LLM benchmarks. Additionally, we explain our modeling choices, training process, and evaluation methodology. We release Training Chronicles (Appendix C) detailing our experience in training BloombergGPT.
Key Findings
1
BloombergGPT is a 50 billion parameter large language model specifically trained for the financial domain.
2
BloombergGPT maintains competitive performance on standard general LLM benchmarks while improving financial task performance.
3
Mixed-domain training yields a model that outperforms existing models on financial tasks by significant margins.
4
The paper documents modeling choices, training process, evaluation methodology, and provides detailed Training Chronicles (Appendix C).
5
Training corpus comprises a 363 billion token finance-specific dataset plus 345 billion tokens from general-purpose datasets.
Research Object
BloombergGPT, a 50-billion-parameter large language model trained on financial and general-purpose text
Research Subject
Performance and behavior of BloombergGPT on financial NLP tasks and general LLM benchmarks as influenced by mixed-domain (financial + general) large-scale training data and training choices
Publication Details
Publication Date
2023-03-30
Journal
Publisher
ISSN
Cited by
309
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by4
A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly2024
Recommender Systems in the Era of Large Language Models (LLMs)2024
Financial Sentiment Analysis: Techniques and Applications2024
RenAIssance: A Survey Into AI Text-to-Image Generation in the Era of Large Model2024