1 min readExecutive Guide

Executive Guide

The Limits of Automatic Evaluation of Creativity in Large Language Models

Author
Aziz Shuaib Ausi
Published
28 August 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

Research investigating the evaluation of creativity in Large Language Models (LLMs) has found significant discrepancies between automatic evaluation methods and human judgment. Human assessments of both human- and AI-generated content reveal that current automated metrics, including LLM-as-a-Judge approaches, do not reliably capture human perceptions of creativity.

Checking access…

Research investigating the evaluation of creativity in Large Language Models (LLMs) has found significant discrepancies between automatic evaluation methods and human judgment. Human assessments of both human- and AI-generated content reveal that current automated metrics, including LLM-as-a-Judge approaches, do not reliably capture human perceptions of creativity.

Why it matters

This research is strategically important as it highlights a fundamental limitation in the current development and deployment of AI systems intended for creative tasks. Organizations relying on automated methods to assess creative outputs from LLMs may be operating on flawed metrics, potentially misjudging the quality, originality, or value of AI-generated content. This misalignment necessitates a re-evaluation of how creative AI applications are developed, tested, and integrated into workflows.

Key insights

  • LLMs are capable of generating text that challenges human performance in creative domains.
  • Evaluating creativity in LLM-generated content remains a significant challenge.
  • Current automatic evaluation methods do not reliably align with human judgments of creativity.
  • The study compared human evaluations across 11 creativity dimensions with automated objective metrics and LLM-as-a-Judge evaluations.
  • Experiments revealed substantial misalignment between automatic evaluations and human assessments.
  • LLM-based judges exhibit a systematic preference for certain characteristics, suggesting bias in their evaluation.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.23705

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). The Limits of Automatic Evaluation of Creativity in Large Language Models. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00740

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00740
Version
v1.0 · r0
Issued
28 August 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication