Executive Guide
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks
- Author
- Aziz Shuaib Ausi
- Published
- August 19, 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
Research on legal multiple-choice benchmarks indicates that certain AI models can achieve above-chance scores by solely analyzing the provided options, without needing the actual question. This suggests a potential flaw in benchmark design or a sophisticated pattern recognition capability within models, where the options themselves contain sufficient information for correct selection. The study focused on UA-JudgeExam, a legal qualification test, revealing that a significant portion of items could be answered blind.
Research on legal multiple-choice benchmarks indicates that certain AI models can achieve above-chance scores by solely analyzing the provided options, without needing the actual question. This suggests a potential flaw in benchmark design or a sophisticated pattern recognition capability within models, where the options themselves contain sufficient information for correct selection. The study focused on UA-JudgeExam, a legal qualification test, revealing that a significant portion of items could be answered blind.
Why it matters
This research highlights a critical vulnerability in the design and evaluation of benchmarks, particularly in high-stakes domains like legal assessment. The ability of AI models to exploit structural or content-based cues in options, independent of questions, undermines the validity of current assessment methodologies and calls into question the genuine understanding models possess.
Key insights
- AI models can score above chance on multiple-choice legal benchmarks by only analyzing options, without the question.
- A specific model (Claude Haiku 4.5) achieved a 0.383 score against chance on the UA-JudgeExam when only shown options.
- The 'leak' of information from options is concentrated, with 11.8% of items answered correctly blind across all option orders, significantly higher than chance expectation (0.2 items).
- The ability to answer blind is not due to direct quotation from legislation; search over Ukrainian legislation recovered only a 0.128 score.
- The UA-JudgeExam benchmark consists of 11,990 four-option items from Ukraine's Higher Qualification Commission of Judges.
- Gating out identified 'leak' items retains 8,128 items for more robust evaluation.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.15428
Related publications
Previous
Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots
Next
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Predicting, Evaluating, and Explaining Top Misinformation Spreaders via Archetypal User Behavior
Executive Guide
Pluralistic Human-Robot Interaction: Designing for Robot Interaction with Diverse Communities
Executive Guide
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Executive Guide
Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots
Executive Guide
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark
Executive Guide
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00404
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00404
- Version
- v1.0 · r0
- Issued
- 8/19/2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.