1 min readExecutive Guide

Executive Guide

Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes

Author
Aziz Shuaib Ausi
Published
28 August 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

A study evaluated the efficacy of GPT-5.5-based multimodal AI in grading handwritten physics assessments, including Olympiad theory and experiment components and a university quantum mechanics examination. The AI graded 10,364 scanned pages from 520 submissions across 416 unique candidates. Initial AI grading was followed by a revised round incorporating feedback from disagreement analysis, with the AI operating without knowledge of human scores or previous AI-human comparisons.

Checking access…

A study evaluated the efficacy of GPT-5.5-based multimodal AI in grading handwritten physics assessments, including Olympiad theory and experiment components and a university quantum mechanics examination. The AI graded 10,364 scanned pages from 520 submissions across 416 unique candidates. Initial AI grading was followed by a revised round incorporating feedback from disagreement analysis, with the AI operating without knowledge of human scores or previous AI-human comparisons.

Why it matters

The capability for large-scale AI grading of complex, handwritten assessments presents a potential paradigm shift in educational and evaluation processes. This technology could significantly enhance efficiency and scalability in high-stakes grading environments, impacting resource allocation and the fairness of selection outcomes.

Key insights

  • Multimodal AI, specifically GPT-5.5, can process and grade handwritten physics solutions.
  • The study covered a large dataset: 10,364 scanned pages from 520 submissions by 416 unique candidates.
  • Assessments included a national Physics Olympiad theory examination, an Olympiad selection camp with theory and experiment, and a university quantum mechanics examination.
  • AI grading was performed twice, with the second round incorporating refined instructions based on initial disagreement analysis.
  • The AI operated independently, without access to official human marks or prior AI-human comparison data during grading.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.20521

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00660

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00660
Version
v1.0 · r0
Issued
28 August 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication