LLM Selection and Vector Database Tuning: A Methodology for Enhancing RAG Systems

Выбор LLM и настройка векторной базы данных: методика улучшения RAG-систем
Łukasz Pawlik
2025-10-10

LLM selectionQdrant vector databasechunk size / context window trade-offsretrieval-augmented generation (RAG)vector embedding models
With the increasing popularity of large language models (LLMs), retrieval-augmented generation (RAG) systems are gaining importance, enabling the use of internal company data to generate precise and relevant responses. The aim of this study was to develop a comprehensive methodology for measuring and optimizing RAG systems, focusing on analyzing the impact of key parameters such as chunk size, vector embedding models, and LLM selection on system effectiveness. Experiments were conducted on a RAG system using a large biographical dataset, stored in a Qdrant vector database, allowing for in-depth analysis in the context of long text data. The results indicated that optimizing RAG systems necessitates considering various factors, including LLM context window size, computational power, and processing costs. The selection of optimal parameters and LLM is a trade-off between response quality, computational cost, and hardware limitations. This study provides practical guidance for engineers and researchers working on improving RAG-based systems, enabling informed decisions regarding RAG system configuration in various business contexts.
1
Developed a comprehensive methodology for measuring and optimizing RAG systems focusing on chunk size, vector embedding models, and LLM selection.
2
Experiments used a large biographical dataset stored in a Qdrant vector database to analyze RAG behavior with long text data.
3
Optimal parameter and LLM selection requires trade-offs among response quality, computational cost, and hardware limitations.
4
Provides practical guidance enabling informed RAG configuration decisions for engineers and researchers in business contexts.
5
RAG system effectiveness depends on LLM context window size, computational power, and processing costs.

Retrieval-augmented generation (RAG) system built on a Qdrant vector database using a large biographical dataset

Effects of parameter choices (chunk size, vector embedding model, LLM selection, context window size) and resource constraints (computational power, processing costs, hardware limits) on RAG system effectiveness and trade-offs between response quality and cost

Publication Details
Publication Date
2025-10-10
Journal
Publisher
ISSN
Cited by
3
Access Type
Author Information
Authors
Łukasz Pawlik
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%