Structured information extraction from scientific text with large language models

Структурированное извлечение информации из научных текстов с помощью больших языковых моделей
Kristin A. Persson, Gerbrand Ceder, John Dagdelen, Anubhav Jain, Nicholas Walker, Sang‐Hoon Lee, Alexander Dunn, Andrew Rosen
2024-02-15

joint named entity recognition and relation extractionlarge language modelsmaterials chemistrymetal-organic frameworksstructured information extraction
Extracting structured knowledge from scientific text remains a challenging task for machine learning models. Here, we present a simple approach to joint named entity recognition and relation extraction and demonstrate how pretrained large language models (GPT-3, Llama-2) can be fine-tuned to extract useful records of complex scientific knowledge. We test three representative tasks in materials chemistry: linking dopants and host materials, cataloging metal-organic frameworks, and general composition/phase/morphology/application information extraction. Records are extracted from single sentences or entire paragraphs, and the output can be returned as simple English sentences or a more structured format such as a list of JSON objects. This approach represents a simple, accessible, and highly flexible route to obtaining large databases of structured specialized scientific knowledge extracted from research papers.
1
A simple joint named-entity recognition and relation-extraction approach enables large language models to extract structured scientific knowledge.
2
Fine-tuned GPT-3 and Llama-2 models successfully extract complex materials-chemistry records from scientific text.
3
Knowledge can be extracted from individual sentences or entire paragraphs and returned as English sentences or structured JSON objects.
4
The approach handles dopant–host linking, metal-organic framework cataloging, and composition, phase, morphology, and application extraction.
5
The method provides an accessible and flexible route for building large databases of specialized scientific knowledge from research papers.

Pretrained large language models fine-tuned to perform joint named entity recognition and relation extraction on scientific text

joint named entity and relation extraction of complex materials-chemistry information from scientific text using fine-tuned large language models

Publication Details
Publication Date
2024-02-15
Journal
Publisher
ISSN
Cited by
707
Access Type
Author Information
Authors
Kristin A. Persson
Gerbrand Ceder
John Dagdelen
Anubhav Jain
Nicholas Walker
Sang‐Hoon Lee
Alexander Dunn
Andrew Rosen
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%