Knowledge Resource · Open access
Research Summary: Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
- Original authors
- Attribution requires verification
- Original source
- arXiv — Computers and Society
- Summary & Analysis prepared by
- Aziz Shuaib Ausi
- Resource type
- Research Summary / Knowledge Resource
- Resource published on AZIZ OS
- 6 October 2026
- Reading time
- 1 min
- Publication type
- Knowledge Resource
- Availability
- Open access
About this Summary & Analysis
AZIZ OS provides independently prepared summaries and analytical interpretations of externally published research and knowledge sources. The underlying works remain attributable to their original authors and rights holders. This resource is intended to improve accessibility and understanding and does not replace the original publication.
Research indicates that while large language model (LLM) judges can consistently rank AI outputs against workplace requirements, they demonstrate significant variability and inaccuracy in determining acceptance rates. A new audit suite, O*NET-BENCH, reveals that despite high agreement on ranking, LLM judges show a broad range of acceptable response estimations (3.0%-97.9%) compared to human worker assessments (61.1%). This suggests that current LLM judge configurations are unreliable for quantifying occupational AI performance, despite their ability to order responses by quality.
Why it matters
The widespread use of AI in occupational settings necessitates reliable evaluation methods. This research highlights a critical discrepancy: while LLMs can order outcomes by quality, their ability to accurately quantify performance or acceptance rates is highly variable and often misaligned with human benchmarks. This divergence poses a significant risk to the credibility and utility of AI systems intended for workplace integration, especially when these systems are used to make decisions about human-level performance or quality standards.
Key insights
- LLM judges are increasingly deployed to evaluate AI outputs against occupational standards.
- The study introduces O*NET-BENCH, an audit suite based on 45,796 worker ratings, to assess LLM judge performance.
- 33 LLM judge configurations across six model families were evaluated on 4,501 test ratings.
- 25 configurations achieved a tie-aware pair accuracy of at least 0.60, indicating reasonable agreement on ranking.
- A train-fitted response-only TF-IDF baseline nearly matched the strongest LLM judge configurations in ranking ability.
- LLM judges estimated acceptance rates ranging from 3.0% to 97.9%, a stark contrast to human worker assessments of 61.1%.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2610.02492
Related resources
The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
Knowledge Resource
A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential
Knowledge Resource
"Lighting The Way For Those Not Here": How Can Technology Researchers Help Resist the Missing and Murdered Indigenous Relatives (MMIR) Crisis?
Knowledge Resource
Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
Knowledge Resource
Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact
Knowledge Resource
Social bot detection in the age of ChatGPT: Challenges and opportunities
Knowledge Resource
Citation
Cite the original work (APA 7)
The original source is authoritative for this citation. Cite the source publication directly — this attribution is pending verification. Open the original source.
Verification
This is an authenticated AZIZ OS resource record.
- Verification ID
- ASA-EXE-2026-01240
- Version
- v1.0 · r0
- Issued
- 6 October 2026
- Resource prepared by
- Aziz Shuaib Ausi
- Resource status
- Research Summary / Knowledge Resource
- Underlying work
- Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
- Original authors
- Attribution requires verification
- Original source
- arXiv — Computers and Society
- Provenance status
- Attribution requires verification
- Rights
- Underlying publication rights remain with the respective copyright holder(s). Refer to the original source for authoritative publication and licensing information.
This verification confirms the AZIZ OS resource record and its documented provenance. It does not establish authorship of the underlying external work.