Advancing AI in Higher Education: A Comparative Study of Large Language Model-Based Agents for Exam Question Generation, Improvement, and Evaluation

Развитие искусственного интеллекта в высшем образовании: сравнительное исследование агентов на основе больших языковых моделей для генерации, совершенствования и оценки экзаменационных вопросов
Vlatko Nikolovski, Dimitar Trajanov, Ivan Chorbev
2025-03-04

Bloom’s taxonomyVectorGraphRAGexam question generationlarge language modelsmixed-effects modeling
The transformative capabilities of large language models (LLMs) are reshaping educational assessment and question design in higher education. This study proposes a systematic framework for leveraging LLMs to enhance question-centric tasks: aligning exam questions with course objectives, improving clarity and difficulty, and generating new items guided by learning goals. The research spans four university courses—two theory-focused and two application-focused—covering diverse cognitive levels according to Bloom’s taxonomy. A balanced dataset ensures representation of question categories and structures. Three LLM-based agents—VectorRAG, VectorGraphRAG, and a fine-tuned LLM—are developed and evaluated against a meta-evaluator, supervised by human experts, to assess alignment accuracy and explanation quality. Robust analytical methods, including mixed-effects modeling, yield actionable insights for integrating generative AI into university assessment processes. Beyond exam-specific applications, this methodology provides a foundational approach for the broader adoption of AI in post-secondary education, emphasizing fairness, contextual relevance, and collaboration. The findings offer a comprehensive framework for aligning AI-generated content with learning objectives, detailing effective integration strategies, and addressing challenges such as bias and contextual limitations. Overall, this work underscores the potential of generative AI to enhance educational assessment while identifying pathways for responsible implementation.
1
Mixed-effects modeling provides evidence-based insights into integrating generative AI into university assessment processes across varied question categories and structures.
2
The evaluation covers four university courses, including theory- and application-focused subjects and diverse cognitive levels based on Bloom’s taxonomy.
3
The framework supports responsible adoption of AI in higher education by emphasizing fairness, contextual relevance, collaboration, and mitigation of bias and contextual limitations.
4
The study develops a systematic LLM-based framework for aligning exam questions with course objectives, improving clarity and difficulty, and generating new learning-goal-driven items.
5
VectorRAG, VectorGraphRAG, and a fine-tuned LLM are compared using a human-supervised meta-evaluator that assesses alignment accuracy and explanation quality.

LLM-based agents for university exam question generation, improvement, and evaluation

their accuracy, explanation quality, and alignment of exam questions with course objectives and learning goals across diverse cognitive levels

Publication Details
Publication Date
2025-03-04
Journal
Publisher
ISSN
Cited by
36
Access Type
Author Information
Authors
Vlatko Nikolovski
Dimitar Trajanov
Ivan Chorbev
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%