Document Type

Conference Paper

Publication Date

2026

DOI

10.55549/epess.1037

Publication Title

The Eurasia Proceedings of Educational and Social Sciences

Volume

49

Pages

228-244

Conference Name

ICRET 2026: International Conference on Research in Education and Technology, April 2-5, 2026, Konya, Türkiye

Abstract

Automated essay evaluation using large language models (LLMs) has emerged as a promising approach to support scalable and consistent educational assessment. However, the effectiveness of LLM-based grading varies significantly across evaluation dimensions and is highly influenced by prompt design and model selection. In this study, we evaluate five state-of-the-art LLMs across five rubric-based categories: Relevance to Question, Reasoning and Critical Thinking, Evidence and Examples, Organization, and Clarity and Writing Quality. We systematically investigate the impact of three prompting strategies, including rubric-only prompting, exemplar-based prompting (with and without rubric guidance)(Original and Refined prompt designs) incorporating structured instructions. Additionally, a prompt ablation study is conducted to analyze the effect of instruction detail and reasoning guidance on grading accuracy. Our results demonstrate that no single model consistently outperforms others across all categories. Instead, each LLM exhibits strengths in specific evaluation dimensions, motivating a category-wise model selection approach. Furthermore, refined and structured prompting strategies significantly improve evaluation performance, with chain-of-thought-style instructions yielding the highest accuracy. These findings highlight the importance of both prompt engineering and task-specialized model allocation in developing robust LLM-based grading systems. The study provides evidence supporting multi-agent frameworks, where different LLMs can be assigned to distinct evaluation tasks to enhance overall grading reliability and alignment with human judgments.

Rights

 © 2026 Published by ISRES Publishing.

This is an Open Access article distributed under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) License, permitting all non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

Original Publication Citation

Yalpi, D., Kasula, S. K., & Mukkamala, R. (2026). Guidelines for automatic grading of student essays using large language models. The Eurasia Proceedings of Educational and Social Sciences, 49, 228-244. https://doi.org/10.55549/epess.1037 

ORCID

0009-0007-2205-9723 (Yalpi), 0009-0001-9024-1873 (Kasula), 0000-0001-6323-9789 (Mukkamala)

Share

COinS