Executive Guide
Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
- Author
- Aziz Shuaib Ausi
- Published
- 28 August 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
A study evaluated the efficacy of GPT-5.5-based multimodal AI in grading handwritten physics assessments, including Olympiad theory and experiment components and a university quantum mechanics examination. The AI graded 10,364 scanned pages from 520 submissions across 416 unique candidates. Initial AI grading was followed by a revised round incorporating feedback from disagreement analysis, with the AI operating without knowledge of human scores or previous AI-human comparisons.
A study evaluated the efficacy of GPT-5.5-based multimodal AI in grading handwritten physics assessments, including Olympiad theory and experiment components and a university quantum mechanics examination. The AI graded 10,364 scanned pages from 520 submissions across 416 unique candidates. Initial AI grading was followed by a revised round incorporating feedback from disagreement analysis, with the AI operating without knowledge of human scores or previous AI-human comparisons.
Why it matters
The capability for large-scale AI grading of complex, handwritten assessments presents a potential paradigm shift in educational and evaluation processes. This technology could significantly enhance efficiency and scalability in high-stakes grading environments, impacting resource allocation and the fairness of selection outcomes.
Key insights
- Multimodal AI, specifically GPT-5.5, can process and grade handwritten physics solutions.
- The study covered a large dataset: 10,364 scanned pages from 520 submissions by 416 unique candidates.
- Assessments included a national Physics Olympiad theory examination, an Olympiad selection camp with theory and experiment, and a university quantum mechanics examination.
- AI grading was performed twice, with the second round incorporating refined instructions based on initial disagreement analysis.
- The AI operated independently, without access to official human marks or prior AI-human comparison data during grading.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.20521
Related publications
Previous
Interaction Effects Between Learner Characteristics and Dialogue Format in TTS Dialogue-Based Lessons
Next
The Substitution Escrow Threshold: When "Compatible With" Becomes Safe Enough to Buy
Embedding inter- and transdisciplinary sustainability skills and knowledge development in higher education: perspectives from an innovative new degree
Executive Guide
Critical thinking as a predictor of task functionality and artificial intelligence use among university students. A PLS-SEM approach
Executive Guide
Cognitive emotion regulation as a statistical mediator of the association between autistic traits and academic performance in university students
Executive Guide
AI self-efficacy as a predictor of satisfaction with studies: the mediating role of research motivation among Peruvian University students
Executive Guide
Generative AI and linguistic creativity in digitally multilingual higher education
Executive Guide
Digital teaching and learning strategies for enhancing self-directed learning in remote ODeL environments: evidence from Zimbabwe Open University
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00660
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00660
- Version
- v1.0 · r0
- Issued
- 28 August 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.