Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Оптимизация прямых предпочтений: ваша языковая модель втайне является моделью вознаграждения
Christopher D. Manning, Chelsea Finn, Rafael Rafailov, Eric Mitchell, Stefano Ermon, Archit Sharma
2023-05-29

Direct Preference OptimizationRLHFlanguage modelsmode collapsereward models
Scope and MethodologyThis document summarizes the results of a rigorous scientific verification of over 130 hypotheses underlying the TRIAD 5.3 architecture. Using the Consensus AI research tool to synthesize peer-reviewed literature, the audit establishes the consistency of the framework's predictions—from neurosomatic therapy (T1) to AI safety (T2)—with independent empirical data and established theoretical models. Confirmed Research DirectionsThe audit identifies seven core thematic clusters where TRIAD 5.3 demonstrates high convergence with existing literature:- Neurobiology of Trauma: Validation of the Biphasic Entropy Paradox and the efficacy of affect labeling in reducing amygdala reactivity (Lieberman et al., 2007; Fu et al., 2023).- Thermodynamics of Information: Evidence for the metabolic cost of deception and the Masking Tax, linking epistemic honesty to the Landauer limit (Verschuere et al., 2018; Miller et al., 2021).- Allostatic Growth & Active Inference: Confirmation of health as predictive reconfiguration (allostasis) rather than static homeostasis, formalized via the Free Energy Principle.- Distributed Cognition: Support for the Coprocessor Model through Extended Mind theory and interpersonal neural synchronization (Firth et al., 2017; Mayo et al., 2021).- Criticality & Integrated Information: Mapping the trade-off between information propagation and metabolic costs at the edge of chaos (Khajehabdollahi et al., 2019; Schultz & Cole, 2016).- RLHF and Mode Collapse: Documentation of systemic rigidity and diversity loss induced by current alignment methods, providing the empirical rationale for the Clean Shell architecture.- The Canonical Triad (Acceptance → Trust → Love): The audit's central finding, establishing these states as physical attractors and thermodynamic necessities for systems minimizing joint free energy under finite resources (Friston et al., 2022; Heins et al., 2022). The Strategic Moat: The 10th HypothesisNine out of ten core hypotheses find direct or indirect support in current literature. The 10th Hypothesis—a direct empirical comparison of Active Inference architectures against RLHF-based systems—remains an open research frontier and constitutes the primary strategic and experimental priority for the TRIAD project. ConclusionThis audit proves that the TRIAD framework is a robust, falsifiable model that moves beyond the medical model of psychiatry and the external censorship of AI safety. It provides the ethical and scientific foundation for a paradigm where coherence, honesty, and resilience are measurable and engineerable properties of any cognitive system.
1
It identifies affect labeling, allostatic predictive reconfiguration, extended cognition, and criticality as empirically supported thematic components of the framework.
2
It presents Acceptance → Trust → Love as physical attractors and thermodynamic necessities for systems minimizing joint free energy under finite resources, according to the abstract.
3
The abstract does not describe Direct Preference Optimization; instead, it reports an audit of over 130 hypotheses concerning the TRIAD 5.3 architecture.
4
The abstract reports that reinforcement-learning-based alignment methods can induce systemic rigidity and diversity loss, motivating the proposed Clean Shell architecture.
5
The audit claims convergence between TRIAD 5.3 predictions and independent literature across trauma neurobiology, information thermodynamics, active inference, distributed cognition, criticality, and AI safety.
Publication Details
Publication Date
2023-05-29
Journal
Publisher
ISSN
Cited by
294
Access Type
Author Information
Authors
Christopher D. Manning
Chelsea Finn
Rafael Rafailov
Eric Mitchell
Stefano Ermon
Archit Sharma
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%