1 min readKnowledge Resource

Knowledge Resource

LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

Author
Aziz Shuaib Ausi
Published
7 September 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
Checking access…

Research auditing the use of Large Language Models (LLMs) as essay graders in educational measurement has identified significant issues beyond simple agreement statistics. The study, conducted across two languages and multiple LLM providers and versions, found substantial variability in judge severity, the presence of halo effects, and concerning instability across different LLM versions. These findings suggest that LLMs, when used for scoring, exhibit characteristics analogous to human rater biases and inconsistencies, which are critical considerations for their deployment in high-stakes assessment contexts.

Why it matters

The increasing integration of Large Language Models into automated assessment systems necessitates a thorough understanding of their reliability and potential biases. These findings highlight critical risks associated with deploying LLMs for high-stakes grading, potentially undermining the fairness and validity of evaluations. Addressing these inconsistencies is paramount for maintaining public trust and ensuring equitable outcomes in educational and professional assessment landscapes.

Key insights

  • LLM judges exhibit considerable variation in severity, with differences spanning hundreds of points on a 0-1000 scale in one dataset and 15-33% of the score range in another.
  • Halo effects, where a general impression of an essay influences specific scoring dimensions, are present in LLM judgments.
  • Significant shifts and instability were observed across different versions of LLMs, indicating that scores generated by these models are not consistent over time or model updates.
  • The study treated LLM judges as human-like raters, applying established educational measurement techniques like many-facet Rasch severity, residual halo analysis, and generalizability studies.
  • Analysis was conducted on public corpora in English and Portuguese, using 2,377 essays, 12 judges, 4 providers, and 5 version contrasts, yielding a comprehensive score tensor.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.29517

Citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00165

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXE-2026-00165
Version
v1.0 · r0
Issued
7 September 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication