Knowledge Resource
"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
- Author
- Aziz Shuaib Ausi
- Published
- 11 September 2026
- Reading time
- 1 min
- Publication type
- Knowledge Resource
- Availability
- Open access
Recent research from arXiv reveals a significant methodological issue, termed "Mirror" contamination, in the evaluation of Large Language Models (LLMs) for predicting depression scores. Studies that use LLM responses derived directly from depression assessment language to predict scores on those same assessments yield near-perfect, but potentially misleading, prediction results. While non-mirror evaluations, using life history interviews, also showed high prediction capabilities, this indicates a need for more robust LLM evaluation methodologies in health-related applications.
Why it matters
This research highlights critical challenges in validating AI systems, particularly LLMs, for sensitive applications like health assessments. Understanding and mitigating methodological biases, such as criterion contamination, is crucial for developing trustworthy and ethically sound AI solutions and ensuring their reliable deployment in real-world scenarios. It underscores the necessity for rigorous and varied testing protocols to accurately gauge AI performance and avoid overstating capabilities.
Key insights
- LLM evaluations where language responses are directly tied to the assessment being predicted ('Mirror' evaluations) show near-perfect prediction of depression scores, indicating criterion contamination.
- A study with 110 participants found that 'Mirror' evaluations were near-perfect when LLMs predicted depression scores from structured diagnostic interviews.
- 'Non-Mirror' evaluations, using life history interviews, also demonstrated prediction sizes considered outstanding in psychology.
- Both 'Mirror' and 'Non-Mirror' LLM predictions correlated with Patient Health Questionnaire (PHQ-9) scores, suggesting some underlying predictive ability beyond contamination.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2508.05830
Related intelligence and resources
Previous
Corporate report: Academy trusts: findings from DfE's assurance work
Next
A global mobile network coverage raster product at 1km resolution, 1999--2030
INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives
Knowledge Resource
Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation
Knowledge Resource
Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms
Knowledge Resource
A global mobile network coverage raster product at 1km resolution, 1999--2030
Knowledge Resource
Corporate report: Academy trusts: findings from DfE's assurance work
Knowledge Resource
Introducing the Agents API
Knowledge Resource
Citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00415
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXE-2026-00415
- Version
- v1.0 · r0
- Issued
- 11 September 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.