Knowledge Resource
LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
- Author
- Aziz Shuaib Ausi
- Published
- 7 September 2026
- Reading time
- 1 min
- Publication type
- Knowledge Resource
- Availability
- Open access
Research auditing the use of Large Language Models (LLMs) as essay graders in educational measurement has identified significant issues beyond simple agreement statistics. The study, conducted across two languages and multiple LLM providers and versions, found substantial variability in judge severity, the presence of halo effects, and concerning instability across different LLM versions. These findings suggest that LLMs, when used for scoring, exhibit characteristics analogous to human rater biases and inconsistencies, which are critical considerations for their deployment in high-stakes assessment contexts.
Why it matters
The increasing integration of Large Language Models into automated assessment systems necessitates a thorough understanding of their reliability and potential biases. These findings highlight critical risks associated with deploying LLMs for high-stakes grading, potentially undermining the fairness and validity of evaluations. Addressing these inconsistencies is paramount for maintaining public trust and ensuring equitable outcomes in educational and professional assessment landscapes.
Key insights
- LLM judges exhibit considerable variation in severity, with differences spanning hundreds of points on a 0-1000 scale in one dataset and 15-33% of the score range in another.
- Halo effects, where a general impression of an essay influences specific scoring dimensions, are present in LLM judgments.
- Significant shifts and instability were observed across different versions of LLMs, indicating that scores generated by these models are not consistent over time or model updates.
- The study treated LLM judges as human-like raters, applying established educational measurement techniques like many-facet Rasch severity, residual halo analysis, and generalizability studies.
- Analysis was conducted on public corpora in English and Portuguese, using 2,377 essays, 12 judges, 4 providers, and 5 version contrasts, yielding a comprehensive score tensor.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.29517
Related intelligence and resources
Previous
The Landscape of Generative AI in Information Systems: A Synthesis of Secondary Reviews and Research Agendas
Next
Shifting from Injection to Interaction: Rethinking Web Security in the Age of LLMs and Beyond
WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing
Knowledge Resource
Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs
Knowledge Resource
Bridging Formal and Perceived Fairness: Development of an Interdisciplinary Framework in Algorithmic Decision-Making
Knowledge Resource
CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling
Knowledge Resource
Affective publics in Arabic YouTube
Knowledge Resource
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
Knowledge Resource
Citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00165
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXE-2026-00165
- Version
- v1.0 · r0
- Issued
- 7 September 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.