Advancing AI in Higher Education: A Comparative Study of Large Language Model-Based Agents for Exam Question Generation, Improvement, and Evaluation
Развитие искусственного интеллекта в высшем образовании: сравнительное исследование агентов на основе больших языковых моделей для генерации, совершенствования и оценки экзаменационных вопросов
2025-03-04
SCID: 54.1/w2s8jcm7
Discuss with AI
Bloom’s taxonomyVectorGraphRAGexam question generationlarge language modelsmixed-effects modeling
Figures from the paper
Abstract (AI)
The transformative capabilities of large language models (LLMs) are reshaping educational assessment and question design in higher education. This study proposes a systematic framework for leveraging LLMs to enhance question-centric tasks: aligning exam questions with course objectives, improving clarity and difficulty, and generating new items guided by learning goals. The research spans four university courses—two theory-focused and two application-focused—covering diverse cognitive levels according to Bloom’s taxonomy. A balanced dataset ensures representation of question categories and structures. Three LLM-based agents—VectorRAG, VectorGraphRAG, and a fine-tuned LLM—are developed and evaluated against a meta-evaluator, supervised by human experts, to assess alignment accuracy and explanation quality. Robust analytical methods, including mixed-effects modeling, yield actionable insights for integrating generative AI into university assessment processes. Beyond exam-specific applications, this methodology provides a foundational approach for the broader adoption of AI in post-secondary education, emphasizing fairness, contextual relevance, and collaboration. The findings offer a comprehensive framework for aligning AI-generated content with learning objectives, detailing effective integration strategies, and addressing challenges such as bias and contextual limitations. Overall, this work underscores the potential of generative AI to enhance educational assessment while identifying pathways for responsible implementation.
Key Findings
1
Mixed-effects modeling provides evidence-based insights into integrating generative AI into university assessment processes across varied question categories and structures.
2
The evaluation covers four university courses, including theory- and application-focused subjects and diverse cognitive levels based on Bloom’s taxonomy.
3
The framework supports responsible adoption of AI in higher education by emphasizing fairness, contextual relevance, collaboration, and mitigation of bias and contextual limitations.
4
The study develops a systematic LLM-based framework for aligning exam questions with course objectives, improving clarity and difficulty, and generating new learning-goal-driven items.
5
VectorRAG, VectorGraphRAG, and a fine-tuned LLM are compared using a human-supervised meta-evaluator that assesses alignment accuracy and explanation quality.
Research Object
LLM-based agents for university exam question generation, improvement, and evaluation
Research Subject
their accuracy, explanation quality, and alignment of exam questions with course objectives and learning goals across diverse cognitive levels
Publication Details
Publication Date
2025-03-04
Journal
Publisher
ISSN
Cited by
36
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest