ai
LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
arXiv: Computers and SocietyInternationalHigh confidence1 min
What changed
Research auditing the use of Large Language Models (LLMs) as essay graders in educational measurement has identified significant issues beyond simple agreement statistics. The study, conducted across two languages and multiple LLM providers and versions, found substantial variability in judge severity, the presence of halo effects, and concerning instability across different LLM versions. These findings suggest that LLMs, when used for scoring, exhibit characteristics analogous to human rater biases and inconsistencies, which are critical considerations for their deployment in high-stakes assessment contexts.
Why it matters
The increasing integration of Large Language Models into automated assessment systems necessitates a thorough understanding of their reliability and potential biases. These findings highlight critical risks associated with deploying LLMs for high-stakes grading, potentially undermining the fairness and validity of evaluations. Addressing these inconsistencies is paramount for maintaining public trust and ensuring equitable outcomes in educational and professional assessment landscapes.
What to watch
LLM judges exhibit considerable variation in severity, with differences spanning hundreds of points on a 0-1000 scale in one dataset and 15-33% of the score range in another.
Forward consideration, not a verified fact.
Reported by arXiv: Computers and Society, International. The document itself is not reproduced here.
Read the original publication