Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
Управление автомобилем с помощью больших языковых моделей: объединение векторной модальности на уровне объектов для объяснимого автономного вождения
2024-05-13
SCID: 54.1/yhnbvh5j
Discuss with AI
Driving QAbehavioral cloningexplainable autonomous drivingobject-level multimodal LLMvectorized numeric modalities
Figures from the paper
Abstract (AI)
Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique objectlevel multimodal LLM architecture that merges vectorized numeric modalities with a pre-trained LLM to improve context understanding in driving situations. We also present a new dataset of 160k QA pairs derived from 10k driving scenarios, paired with high quality control commands collected with RL agent and question answer pairs generated by teacher LLM (GPT-3.5). A distinct pretraining strategy is devised to align numeric vector modalities with static LLM representations using vector captioning language data. We also introduce an evaluation metric for Driving QA and demonstrate our LLM-driver’s proficiency in interpreting driving scenarios, answering questions, and decision-making. Our findings highlight the potential of LLM-based driving action generation in comparison to traditional behavioral cloning. We make our benchmark, datasets, and model available1for further exploration.
Key Findings
1
Develops a vector-captioning pretraining strategy to align numeric vector representations with static pretrained LLM representations.
2
Introduces a dedicated Driving QA evaluation metric and demonstrates proficiency in scenario interpretation, question answering, and driving decision-making.
3
Introduces an object-level multimodal LLM architecture that fuses vectorized numeric driving modalities with a pretrained language model for improved contextual understanding.
4
Releases a dataset containing 160,000 question-answer pairs from 10,000 driving scenarios, alongside reinforcement-learning-derived control commands and GPT-3.5-generated answers.
5
Shows the potential of LLM-based driving action generation compared with traditional behavioral cloning, while releasing the benchmark, datasets, and model for further research.
Research Object
object-level multimodal LLM-based autonomous driving system
Research Subject
context understanding, explainable driving-question answering, and driving action generation from vectorized driving-scene modalities
Publication Details
Publication Date
2024-05-13
Journal
Publisher
ISSN
Cited by
175
Access Type
Author Information
Download PDF
Subscribe to digest