LLM Selection and Vector Database Tuning: A Methodology for Enhancing RAG Systems
Выбор LLM и настройка векторной базы данных: методика улучшения RAG-систем
2025-10-10
SCID: 54.1/b4q63kf7
Discuss with AI
LLM selectionQdrant vector databasechunk size / context window trade-offsretrieval-augmented generation (RAG)vector embedding models
Figures from the paper
Abstract (AI)
With the increasing popularity of large language models (LLMs), retrieval-augmented generation (RAG) systems are gaining importance, enabling the use of internal company data to generate precise and relevant responses. The aim of this study was to develop a comprehensive methodology for measuring and optimizing RAG systems, focusing on analyzing the impact of key parameters such as chunk size, vector embedding models, and LLM selection on system effectiveness. Experiments were conducted on a RAG system using a large biographical dataset, stored in a Qdrant vector database, allowing for in-depth analysis in the context of long text data. The results indicated that optimizing RAG systems necessitates considering various factors, including LLM context window size, computational power, and processing costs. The selection of optimal parameters and LLM is a trade-off between response quality, computational cost, and hardware limitations. This study provides practical guidance for engineers and researchers working on improving RAG-based systems, enabling informed decisions regarding RAG system configuration in various business contexts.
Key Findings
1
Developed a comprehensive methodology for measuring and optimizing RAG systems focusing on chunk size, vector embedding models, and LLM selection.
2
Experiments used a large biographical dataset stored in a Qdrant vector database to analyze RAG behavior with long text data.
3
Optimal parameter and LLM selection requires trade-offs among response quality, computational cost, and hardware limitations.
4
Provides practical guidance enabling informed RAG configuration decisions for engineers and researchers in business contexts.
5
RAG system effectiveness depends on LLM context window size, computational power, and processing costs.
Research Object
Retrieval-augmented generation (RAG) system built on a Qdrant vector database using a large biographical dataset
Research Subject
Effects of parameter choices (chunk size, vector embedding model, LLM selection, context window size) and resource constraints (computational power, processing costs, hardware limits) on RAG system effectiveness and trade-offs between response quality and cost
Publication Details
Publication Date
2025-10-10
Journal
Publisher
ISSN
Cited by
3
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest