Document Type
Conference Paper
Publication Date
2026
DOI
10.55549/epess.1037
Publication Title
The Eurasia Proceedings of Educational and Social Sciences
Volume
49
Pages
228-244
Conference Name
ICRET 2026: International Conference on Research in Education and Technology, April 2-5, 2026, Konya, Türkiye
Abstract
Automated essay evaluation using large language models (LLMs) has emerged as a promising approach to support scalable and consistent educational assessment. However, the effectiveness of LLM-based grading varies significantly across evaluation dimensions and is highly influenced by prompt design and model selection. In this study, we evaluate five state-of-the-art LLMs across five rubric-based categories: Relevance to Question, Reasoning and Critical Thinking, Evidence and Examples, Organization, and Clarity and Writing Quality. We systematically investigate the impact of three prompting strategies, including rubric-only prompting, exemplar-based prompting (with and without rubric guidance)(Original and Refined prompt designs) incorporating structured instructions. Additionally, a prompt ablation study is conducted to analyze the effect of instruction detail and reasoning guidance on grading accuracy. Our results demonstrate that no single model consistently outperforms others across all categories. Instead, each LLM exhibits strengths in specific evaluation dimensions, motivating a category-wise model selection approach. Furthermore, refined and structured prompting strategies significantly improve evaluation performance, with chain-of-thought-style instructions yielding the highest accuracy. These findings highlight the importance of both prompt engineering and task-specialized model allocation in developing robust LLM-based grading systems. The study provides evidence supporting multi-agent frameworks, where different LLMs can be assigned to distinct evaluation tasks to enhance overall grading reliability and alignment with human judgments.
Rights
© 2026 Published by ISRES Publishing.
This is an Open Access article distributed under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) License, permitting all non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Original Publication Citation
Yalpi, D., Kasula, S. K., & Mukkamala, R. (2026). Guidelines for automatic grading of student essays using large language models. The Eurasia Proceedings of Educational and Social Sciences, 49, 228-244. https://doi.org/10.55549/epess.1037
Repository Citation
Yalpi, D., Kasula, S. K., & Mukkamala, R. (2026). Guidelines for automatic grading of student essays using large language models. The Eurasia Proceedings of Educational and Social Sciences, 49, 228-244. https://doi.org/10.55549/epess.1037
ORCID
0009-0007-2205-9723 (Yalpi), 0009-0001-9024-1873 (Kasula), 0000-0001-6323-9789 (Mukkamala)
Included in
Artificial Intelligence and Robotics Commons, Educational Assessment, Evaluation, and Research Commons, Educational Technology Commons, Higher Education Commons