Skip to main content
1 min readKnowledge Resource

Knowledge Resource · Open access

Research Summary: Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

Original authors
Attribution requires verification
Original source
arXiv — Computers and Society
Summary & Analysis prepared by
Aziz Shuaib Ausi
Resource type
Research Summary / Knowledge Resource
Resource published on AZIZ OS
6 October 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
About this Summary & Analysis

AZIZ OS provides independently prepared summaries and analytical interpretations of externally published research and knowledge sources. The underlying works remain attributable to their original authors and rights holders. This resource is intended to improve accessibility and understanding and does not replace the original publication.

Checking access…

Research indicates that while large language model (LLM) judges can consistently rank AI outputs against workplace requirements, they demonstrate significant variability and inaccuracy in determining acceptance rates. A new audit suite, O*NET-BENCH, reveals that despite high agreement on ranking, LLM judges show a broad range of acceptable response estimations (3.0%-97.9%) compared to human worker assessments (61.1%). This suggests that current LLM judge configurations are unreliable for quantifying occupational AI performance, despite their ability to order responses by quality.

Why it matters

The widespread use of AI in occupational settings necessitates reliable evaluation methods. This research highlights a critical discrepancy: while LLMs can order outcomes by quality, their ability to accurately quantify performance or acceptance rates is highly variable and often misaligned with human benchmarks. This divergence poses a significant risk to the credibility and utility of AI systems intended for workplace integration, especially when these systems are used to make decisions about human-level performance or quality standards.

Key insights

  • LLM judges are increasingly deployed to evaluate AI outputs against occupational standards.
  • The study introduces O*NET-BENCH, an audit suite based on 45,796 worker ratings, to assess LLM judge performance.
  • 33 LLM judge configurations across six model families were evaluated on 4,501 test ratings.
  • 25 configurations achieved a tie-aware pair accuracy of at least 0.60, indicating reasonable agreement on ranking.
  • A train-fitted response-only TF-IDF baseline nearly matched the strongest LLM judge configurations in ranking ability.
  • LLM judges estimated acceptance rates ranging from 3.0% to 97.9%, a stark contrast to human worker assessments of 61.1%.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2610.02492

Citation

Cite the original work (APA 7)

The original source is authoritative for this citation. Cite the source publication directly — this attribution is pending verification. Open the original source.

Verification

This is an authenticated AZIZ OS resource record.

Verification ID
ASA-EXE-2026-01240
Version
v1.0 · r0
Issued
6 October 2026
Resource prepared by
Aziz Shuaib Ausi
Resource status
Research Summary / Knowledge Resource
Underlying work
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
Original authors
Attribution requires verification
Original source
arXiv — Computers and Society
Provenance status
Attribution requires verification
Rights
Underlying publication rights remain with the respective copyright holder(s). Refer to the original source for authoritative publication and licensing information.

This verification confirms the AZIZ OS resource record and its documented provenance. It does not establish authorship of the underlying external work.

Verify this resource