Training language models to follow instructions with human feedback
Обучение языковых моделей следованию инструкциям с использованием обратной связи от человека
2022-03-04
SCID: 54.1/8uddase2
Discuss with AI
InstructGPThuman feedbackinstruction followinglanguage model alignmentreinforcement learning from human feedback
Figures from the paper
Abstract (AI)
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.
Key Findings
1
Fine-tuning language models with human feedback provides an approach for aligning outputs with user intent across diverse tasks.
2
Human evaluators preferred outputs from 1.3B-parameter InstructGPT over outputs from 175B-parameter GPT-3, despite 100 times fewer parameters.
3
InstructGPT improves truthfulness and reduces toxic outputs while causing minimal performance regressions on public NLP datasets.
4
InstructGPT still makes simple mistakes, indicating that human-feedback fine-tuning improves alignment but does not eliminate model errors.
5
Supervised fine-tuning on labeler-written demonstrations followed by reinforcement learning from ranked outputs produces the InstructGPT models.
Research Object
InstructGPT language models fine-tuned with human feedback
Research Subject
Alignment with user intent, including instruction-following, truthfulness, toxicity, helpfulness, and task performance
Publication Details
Publication Date
2022-03-04
Journal
Publisher
ISSN
Cited by
4349
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by20
Large language models encode clinical knowledge2023
A Survey on Evaluation of Large Language Models2024
A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly2024
Autonomous chemical research with large language models2023
Large Language Models for Software Engineering: A Systematic Literature Review2024
Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT2023
Augmenting large language models with chemistry tools2024
ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health2023
Generative artificial intelligence2023
Can Large Language Models Transform Computational Social Science?2023
Artificial Intelligence for Predictive Maintenance Applications: Key Components, Trustworthiness, and Future Trends2024
A framework for human evaluation of large language models in healthcare derived from literature review2024
A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges2024
Talking about Large Language Models2024
ChatGPT and Software Testing Education: Promises & Perils2023
A Review of Current Trends, Techniques, and Challenges in Large Language Models (LLMs)2024
A Comprehensive Review on Synergy of Multi-Modal Data and AI Technologies in Medical Diagnosis2024
Vision-language models for medical report generation and visual question answering: a review2024
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways2025
Enhancing Financial Sentiment Analysis via Retrieval Augmented Large Language Models2023