ai
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
arXiv: Computers and SocietyInternationalHigh confidence1 min
What changed
Research indicates that while large language model (LLM) judges can consistently rank AI outputs against workplace requirements, they demonstrate significant variability and inaccuracy in determining acceptance rates. A new audit suite, O*NET-BENCH, reveals that despite high agreement on ranking, LLM judges show a broad range of acceptable response estimations (3.0%-97.9%) compared to human worker assessments (61.1%). This suggests that current LLM judge configurations are unreliable for quantifying occupational AI performance, despite their ability to order responses by quality.
Why it matters
The widespread use of AI in occupational settings necessitates reliable evaluation methods. This research highlights a critical discrepancy: while LLMs can order outcomes by quality, their ability to accurately quantify performance or acceptance rates is highly variable and often misaligned with human benchmarks. This divergence poses a significant risk to the credibility and utility of AI systems intended for workplace integration, especially when these systems are used to make decisions about human-level performance or quality standards.
What to watch
LLM judges are increasingly deployed to evaluate AI outputs against occupational standards.
Forward consideration, not a verified fact.
Reported by arXiv: Computers and Society, International. The document itself is not reproduced here.
Read the original publication