FinBERT : A Large Language Model for Extracting Information from Financial Text*

FinBERT: большая языковая модель для извлечения информации из финансовых текстов
Allen Huang, Hui Wang, Yi Yang
2022-09-30

FinBERTanalyst reportsenvironmental, social, and governance (ESG)financial textsentiment classification
ABSTRACT We develop FinBERT, a state‐of‐the‐art large language model that adapts to the finance domain. We show that FinBERT incorporates finance knowledge and can better summarize contextual information in financial texts. Using a sample of researcher‐labeled sentences from analyst reports, we document that FinBERT substantially outperforms the Loughran and McDonald dictionary and other machine learning algorithms, including naïve Bayes, support vector machine, random forest, convolutional neural network, and long short‐term memory, in sentiment classification. Our results show that FinBERT excels in identifying the positive or negative sentiment of sentences that other algorithms mislabel as neutral, likely because it uses contextual information in financial text. We find that FinBERT's advantage over other algorithms, and Google's original bidirectional encoder representations from transformers model, is especially salient when the training sample size is small and in texts containing financial words not frequently used in general texts. FinBERT also outperforms other models in identifying discussions related to environment, social, and governance issues. Last, we show that other approaches underestimate the textual informativeness of earnings conference calls by at least 18% compared to FinBERT. Our results have implications for academic researchers, investment professionals, and financial market regulators.
1
FinBERT adapts a large language model to finance, incorporating domain knowledge and improving contextual summarization of financial text.
2
FinBERT more accurately identifies positive or negative sentiment in sentences misclassified as neutral by other algorithms, reflecting its use of financial context.
3
FinBERT outperforms other models in detecting environmental, social, and governance discussions and shows that alternative methods underestimate earnings-call informativeness by at least 18%.
4
FinBERT substantially outperforms the Loughran–McDonald dictionary and multiple machine-learning baselines in sentiment classification of analyst-report sentences.
5
FinBERT’s advantage over competing models, including Google’s original BERT, is strongest with small training samples and domain-specific financial terms uncommon in general text.

financial texts, including analyst reports and earnings conference calls

FinBERT’s domain adaptation and performance in extracting contextual sentiment and ESG-related information, particularly under small training samples and in assessing textual informativeness

Publication Details
Publication Date
2022-09-30
Journal
Publisher
ISSN
Cited by
772
Access Type
Author Information
Authors
Allen Huang
Hui Wang
Yi Yang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%