Executive Guide
What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
- Author
- Aziz Shuaib Ausi
- Published
- 28 August 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
A recent study conducted a blind Turing Test to evaluate the performance of leading out-of-the-box Large Language Models (LLMs) on Italian legal professional exams for lawyers, judges, and notaries. The LLMs generated full exam papers, which were anonymized and assessed by expert examiners using real-world criteria. The research revealed significant variations in model performance across different tasks and exams. While some LLMs demonstrated capabilities comparable to or exceeding top human performance in adversarial legal argumentation and doctrinal analysis, all models consistently failed the notary exam, which demands precise legal planning under strict formal and substantive constraints.
A recent study conducted a blind Turing Test to evaluate the performance of leading out-of-the-box Large Language Models (LLMs) on Italian legal professional exams for lawyers, judges, and notaries. The LLMs generated full exam papers, which were anonymized and assessed by expert examiners using real-world criteria. The research revealed significant variations in model performance across different tasks and exams. While some LLMs demonstrated capabilities comparable to or exceeding top human performance in adversarial legal argumentation and doctrinal analysis, all models consistently failed the notary exam, which demands precise legal planning under strict formal and substantive constraints.
Why it matters
This research provides critical insights into the current capabilities and limitations of advanced LLMs in complex professional domains, particularly those requiring precise legal planning and execution. Understanding these performance boundaries is essential for informed decision-making regarding the integration and reliance on AI tools in high-stakes professional environments and for guiding future AI development toward more robust and reliable applications.
Key insights
- Leading LLMs show variable performance across different legal examination tasks.
- Some LLMs achieved human-level or superior performance in adversarial legal argumentation and doctrinal analysis.
- All tested LLMs failed to meet requirements for the notary exam, which requires goal-directed legal planning and adherence to strict constraints.
- The evaluation method used a blind Turing Test with expert examiners and real examination criteria.
- Performance differences were observed both across various models and different types of legal tasks.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.06166
Related publications
Previous
Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus Analysis
Next
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
Non-automatable cognitive skills in higher education in the age of generative AI
Executive Guide
Differential relations of mathematics vocabulary and academic skill performance in middle school
Executive Guide
“Not giving up when I can't express myself”: toward an entrepreneurial mindset framework in English language learning
Executive Guide
AI Can Help One of K-12’s Biggest Challenges: Middle School Reading
Executive Guide
2 Calif. Bills Could Let More Community Colleges Grant 4-Year Degrees
Executive Guide
Alumni Are Forever. Alumni Email Addresses Are Not.
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00698
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00698
- Version
- v1.0 · r0
- Issued
- 28 August 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.