Exploring the Behavior and Performance of Large Language Models: Can LLMs Infer Answers to Questions Involving Restricted Information?

Ángel Cadena-Bautista, Francisco Fernando Lopez-Ponce, Sergio-Luis Ojeda-Trueba, Gerardo Sierra, Gemma Bel-Enguix
2025-01-22

SCID:  54.1/ytd9bqu7
In this paper various LLMs are tested in a specific domain using a Retrieval-Augmented Generation (RAG) system. The study focuses on the performance and behavior of the models and was conducted in Spanish. A questionnaire based on The Bible, which consists of questions that vary in complexity of reasoning, was created in order to evaluate the reasoning capabilities of each model. The RAG system matches a question with the most similar passage from The Bible and feeds the pair to each LLM. The evaluation aims to determine whether each model can reason solely with the provided information or if it disregards the instructions given and makes use of its pretrained knowledge.
Publication Details
Publication Date
2025-01-22
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Ángel Cadena-Bautista
Francisco Fernando Lopez-Ponce
Sergio-Luis Ojeda-Trueba
Gerardo Sierra
Gemma Bel-Enguix
Explore More Research
Use the citation graph to discover related papers and expand your research horizons.
Click any node to explore
Download PDF
100%