Sparks of Artificial General Intelligence: Early experiments with GPT-4
Искры искусственного общего интеллекта: ранние эксперименты с GPT-4
2023-03-22
SCID: 54.1/bnk4bq94
Discuss with AI
GPT-4artificial general intelligencelarge language modelsmultimodal capabilitiesnext-word prediction
Figures from the paper
Abstract (AI)
Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The latest model developed by OpenAI, GPT-4, was trained using an unprecedented scale of compute and data. In this paper, we report on our investigation of an early version of GPT-4, when it was still in active development by OpenAI. We contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models. We discuss the rising capabilities and implications of these models. We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. In our exploration of GPT-4, we put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction. We conclude with reflections on societal influences of the recent technological leap and future research directions.
Key Findings
1
An early version of GPT-4 solved novel, difficult tasks across mathematics, coding, vision, medicine, law, and psychology without specialized prompting.
2
GPT-4 belongs to a newer cohort of large language models exhibiting substantially more general intelligence than previous AI systems.
3
GPT-4’s performance was often close to human-level across diverse domains and frequently surpassed earlier models such as ChatGPT.
4
The breadth and depth of GPT-4’s capabilities led the authors to characterize it as a possible early, incomplete artificial general intelligence system.
5
The study emphasizes GPT-4’s limitations and identifies challenges for developing more comprehensive AGI, potentially requiring paradigms beyond next-word prediction.
Research Object
An early (in-development) version of the GPT-4 large language model
Research Subject
GPT-4's general intelligence capabilities, cross-domain task performance, human-level performance, and limitations
Publication Details
Publication Date
2023-03-22
Journal
Publisher
ISSN
Cited by
1585
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by17
A Survey on Evaluation of Large Language Models2024
The future landscape of large language models in medicine2023
Autonomous chemical research with large language models2023
Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT2023
Generative AI at Work2025
Augmenting large language models with chemistry tools2024
GPT (Generative Pre-Trained Transformer)— A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions2024
Generative artificial intelligence2023
Large Language Models for Software Engineering: Survey and Open Problems2023
Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation2024
Unleashing the potential of prompt engineering for large language models2025
ChatMOF: an artificial intelligence system for predicting and generating metal-organic frameworks using large language models2024
Retrieval augmented generation for large language models in healthcare: A systematic review2025
Diffractive optical computing in free space2024
Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving2024
Human-AI agency in the age of generative AI2025
Evaluating large language models as agents in the clinic2024