ai
What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
arXiv: Computers and SocietyInternationalModerate confidence1 min
What changed
A recent study conducted a blind Turing Test to evaluate the performance of leading out-of-the-box Large Language Models (LLMs) on Italian legal professional exams for lawyers, judges, and notaries. The LLMs generated full exam papers, which were anonymized and assessed by expert examiners using real-world criteria. The research revealed significant variations in model performance across different tasks and exams. While some LLMs demonstrated capabilities comparable to or exceeding top human performance in adversarial legal argumentation and doctrinal analysis, all models consistently failed the notary exam, which demands precise legal planning under strict formal and substantive constraints.
Why it matters
This research provides critical insights into the current capabilities and limitations of advanced LLMs in complex professional domains, particularly those requiring precise legal planning and execution. Understanding these performance boundaries is essential for informed decision-making regarding the integration and reliance on AI tools in high-stakes professional environments and for guiding future AI development toward more robust and reliable applications.
What to watch
Leading LLMs show variable performance across different legal examination tasks.
Forward consideration, not a verified fact.
Reported by arXiv: Computers and Society, International. The document itself is not reproduced here.
Read the original publication