Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

Управление автомобилем с помощью больших языковых моделей: объединение векторной модальности на уровне объектов для объяснимого автономного вождения
Long Chen, Oleg Sinavski, Jan Hünermann, Alice Karnsund, Andrew James Willmott, Danny Birch, Daniel Maund, Jamie Shotton
2024-05-13

Driving QAbehavioral cloningexplainable autonomous drivingobject-level multimodal LLMvectorized numeric modalities
Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique objectlevel multimodal LLM architecture that merges vectorized numeric modalities with a pre-trained LLM to improve context understanding in driving situations. We also present a new dataset of 160k QA pairs derived from 10k driving scenarios, paired with high quality control commands collected with RL agent and question answer pairs generated by teacher LLM (GPT-3.5). A distinct pretraining strategy is devised to align numeric vector modalities with static LLM representations using vector captioning language data. We also introduce an evaluation metric for Driving QA and demonstrate our LLM-driver’s proficiency in interpreting driving scenarios, answering questions, and decision-making. Our findings highlight the potential of LLM-based driving action generation in comparison to traditional behavioral cloning. We make our benchmark, datasets, and model available1for further exploration.
1
Develops a vector-captioning pretraining strategy to align numeric vector representations with static pretrained LLM representations.
2
Introduces a dedicated Driving QA evaluation metric and demonstrates proficiency in scenario interpretation, question answering, and driving decision-making.
3
Introduces an object-level multimodal LLM architecture that fuses vectorized numeric driving modalities with a pretrained language model for improved contextual understanding.
4
Releases a dataset containing 160,000 question-answer pairs from 10,000 driving scenarios, alongside reinforcement-learning-derived control commands and GPT-3.5-generated answers.
5
Shows the potential of LLM-based driving action generation compared with traditional behavioral cloning, while releasing the benchmark, datasets, and model for further research.

object-level multimodal LLM-based autonomous driving system

context understanding, explainable driving-question answering, and driving action generation from vectorized driving-scene modalities

Publication Details
Publication Date
2024-05-13
Journal
Publisher
ISSN
Cited by
175
Access Type
Author Information
Authors
Long Chen
Oleg Sinavski
Jan Hünermann
Alice Karnsund
Andrew James Willmott
Danny Birch
Daniel Maund
Jamie Shotton
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%