1 min readExecutive Guide

Executive Guide

What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

Author
Aziz Shuaib Ausi
Published
28 August 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

A recent study conducted a blind Turing Test to evaluate the performance of leading out-of-the-box Large Language Models (LLMs) on Italian legal professional exams for lawyers, judges, and notaries. The LLMs generated full exam papers, which were anonymized and assessed by expert examiners using real-world criteria. The research revealed significant variations in model performance across different tasks and exams. While some LLMs demonstrated capabilities comparable to or exceeding top human performance in adversarial legal argumentation and doctrinal analysis, all models consistently failed the notary exam, which demands precise legal planning under strict formal and substantive constraints.

Checking access…

A recent study conducted a blind Turing Test to evaluate the performance of leading out-of-the-box Large Language Models (LLMs) on Italian legal professional exams for lawyers, judges, and notaries. The LLMs generated full exam papers, which were anonymized and assessed by expert examiners using real-world criteria. The research revealed significant variations in model performance across different tasks and exams. While some LLMs demonstrated capabilities comparable to or exceeding top human performance in adversarial legal argumentation and doctrinal analysis, all models consistently failed the notary exam, which demands precise legal planning under strict formal and substantive constraints.

Why it matters

This research provides critical insights into the current capabilities and limitations of advanced LLMs in complex professional domains, particularly those requiring precise legal planning and execution. Understanding these performance boundaries is essential for informed decision-making regarding the integration and reliance on AI tools in high-stakes professional environments and for guiding future AI development toward more robust and reliable applications.

Key insights

  • Leading LLMs show variable performance across different legal examination tasks.
  • Some LLMs achieved human-level or superior performance in adversarial legal argumentation and doctrinal analysis.
  • All tested LLMs failed to meet requirements for the notary exam, which requires goal-directed legal planning and adherence to strict constraints.
  • The evaluation method used a blind Turing Test with expert examiners and real examination criteria.
  • Performance differences were observed both across various models and different types of legal tasks.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.06166

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00698

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00698
Version
v1.0 · r0
Issued
28 August 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication