Executive Guide · Open access
Research Summary: How consistent is the algorithm? Examining the intra- and inter-rater reliability of LLM-based writing assessment
- Original authors
- Attribution requires verification
- Original source
- Frontiers in Education
- Summary & Analysis prepared by
- Aziz Shuaib Ausi
- Resource type
- Research Summary / Knowledge Resource
- Resource published on AZIZ OS
- 13 August 2026
- Last updated
- 22 September 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
About this Summary & Analysis
AZIZ OS provides independently prepared summaries and analytical interpretations of externally published research and knowledge sources. The underlying works remain attributable to their original authors and rights holders. This resource is intended to improve accessibility and understanding and does not replace the original publication.
A study published in Frontiers in Education examined the reliability of Large Language Model (LLM)-based writing assessment, specifically using OpenAI's ChatGPT. The research focused on assessing 192 essays from EFL learners, applying a criterion-based analytic rubric. The study found high intra-rater reliability for the LLM, indicating consistency in its scoring over time.
Why it matters
The consistent performance of AI in assessment offers potential for scalable and objective evaluation processes across various domains. Understanding the reliability of LLM-based tools is crucial for their integration into existing frameworks and for developing future strategies that leverage artificial intelligence.
Key insights
- LLM-based scoring demonstrates high intra-rater reliability, meaning consistent scores for the same essays over time.
- OpenAI's ChatGPT was utilized to assess EFL learner essays using a criterion-based analytic rubric.
- The assessment rubric covered task achievement, grammatical range and accuracy, lexical resources, organization, and mechanics.
- The study replicated scoring after three weeks under similar conditions to assess intra-rater reliability.
Source
Frontiers in Education — https://www.frontiersin.org/articles/10.3389/feduc.2026.1861960
Related resources
Previous
Pedagogical prompting rather than technical mastery: Generative AI use by English and English-medium instruction teachers
Next
ESD conceptualization and enactment in Austrian teacher education
Transformative play: integrating outdoor adventure education and the NPI-cycle to facilitate transformative experience
Executive Guide
Cybersecurity Threat Delays Start of Classes at UT San Antonio
Executive Guide
Towards the determination of competencies of the commercial engineer in Chile
Executive Guide
From Atari to EVE Online: Building on 15 Years of AI Research in Games
Executive Guide
Bankrupt Saint Augustine’s Will Not Offer Fall Classes
Executive Guide
Cornell Hopes to Turn Cheating Into Teachable Moment
Executive Guide
Citation
Cite the original work (APA 7)
The original source is authoritative for this citation. Cite the source publication directly — this attribution is pending verification. Open the original source.
Verification
This is an authenticated AZIZ OS resource record.
- Verification ID
- ASA-EXG-2026-00205
- Version
- v1.0 · r0
- Issued
- 13 August 2026
- Resource prepared by
- Aziz Shuaib Ausi
- Resource status
- Research Summary / Knowledge Resource
- Underlying work
- How consistent is the algorithm? Examining the intra- and inter-rater reliability of LLM-based writing assessment
- Original authors
- Attribution requires verification
- Original source
- Frontiers in Education
- Provenance status
- Attribution requires verification
- Rights
- Underlying publication rights remain with the respective copyright holder(s). Refer to the original source for authoritative publication and licensing information.
This verification confirms the AZIZ OS resource record and its documented provenance. It does not establish authorship of the underlying external work.