1 min readKnowledge Resource

Knowledge Resource

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

Author
Aziz Shuaib Ausi
Published
11 September 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
Checking access…

Recent research from arXiv reveals a significant methodological issue, termed "Mirror" contamination, in the evaluation of Large Language Models (LLMs) for predicting depression scores. Studies that use LLM responses derived directly from depression assessment language to predict scores on those same assessments yield near-perfect, but potentially misleading, prediction results. While non-mirror evaluations, using life history interviews, also showed high prediction capabilities, this indicates a need for more robust LLM evaluation methodologies in health-related applications.

Why it matters

This research highlights critical challenges in validating AI systems, particularly LLMs, for sensitive applications like health assessments. Understanding and mitigating methodological biases, such as criterion contamination, is crucial for developing trustworthy and ethically sound AI solutions and ensuring their reliable deployment in real-world scenarios. It underscores the necessity for rigorous and varied testing protocols to accurately gauge AI performance and avoid overstating capabilities.

Key insights

  • LLM evaluations where language responses are directly tied to the assessment being predicted ('Mirror' evaluations) show near-perfect prediction of depression scores, indicating criterion contamination.
  • A study with 110 participants found that 'Mirror' evaluations were near-perfect when LLMs predicted depression scores from structured diagnostic interviews.
  • 'Non-Mirror' evaluations, using life history interviews, also demonstrated prediction sizes considered outstanding in psychology.
  • Both 'Mirror' and 'Non-Mirror' LLM predictions correlated with Patient Health Questionnaire (PHQ-9) scores, suggesting some underlying predictive ability beyond contamination.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2508.05830

Citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00415

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXE-2026-00415
Version
v1.0 · r0
Issued
11 September 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication