Executive Guide
Large language models simulate intersectional synthetic identities with a budget of one to two dimensions
- Author
- Aziz Shuaib Ausi
- Published
- 28 August 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
Research from arXiv highlights that large language models (LLMs) used as synthetic survey respondents struggle to accurately represent intersectional populations. While real-world subgroup opinions become more distinctive with increasing intersectional identities, LLM-simulated respondents exhibit a 'collapse' where a single identity component often explains responses better than additive combinations of multiple identities. This suggests LLMs do not genuinely simulate complex, multi-dimensional identities.
Research from arXiv highlights that large language models (LLMs) used as synthetic survey respondents struggle to accurately represent intersectional populations. While real-world subgroup opinions become more distinctive with increasing intersectional identities, LLM-simulated respondents exhibit a 'collapse' where a single identity component often explains responses better than additive combinations of multiple identities. This suggests LLMs do not genuinely simulate complex, multi-dimensional identities.
Why it matters
This research is strategically important because it reveals a significant limitation in the current capability of large language models to accurately simulate complex human populations, particularly those with intersectional identities. This impacts the reliability and validity of insights derived from synthetic data generated by LLMs, potentially leading to flawed policy, product, or operational decisions if not properly understood and mitigated.
Key insights
- Large language models are increasingly utilized as synthetic survey respondents to access rare intersectional populations cheaply.
- Real-world data from Pew's American Trends Panel shows that subgroup opinion becomes 2.5 times more distinctive as identities intersect.
- LLM-simulated respondents, across eight models and 21 million simulated response distributions, do not replicate this compositional effect.
- For LLM-simulated respondents, a single feature explains a two-feature persona's responses better than an additive combination in 75-82% of subgroups.
- The addition of a third identity feature to LLM-simulated personas provides almost no additional explanatory power for their responses.
- This 'collapse' of intersectional distinctiveness in LLMs persists despite various prompting strategies.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.23005
Related publications
Previous
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds
Next
Self-Reported AI Usage for Learning in Computer Science Education: Relationships with Goal Orientation and Academic Help-Seeking
Embedding inter- and transdisciplinary sustainability skills and knowledge development in higher education: perspectives from an innovative new degree
Executive Guide
Critical thinking as a predictor of task functionality and artificial intelligence use among university students. A PLS-SEM approach
Executive Guide
Cognitive emotion regulation as a statistical mediator of the association between autistic traits and academic performance in university students
Executive Guide
AI self-efficacy as a predictor of satisfaction with studies: the mediating role of research motivation among Peruvian University students
Executive Guide
Generative AI and linguistic creativity in digitally multilingual higher education
Executive Guide
Digital teaching and learning strategies for enhancing self-directed learning in remote ODeL environments: evidence from Zimbabwe Open University
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). Large language models simulate intersectional synthetic identities with a budget of one to two dimensions. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00602
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00602
- Version
- v1.0 · r0
- Issued
- 28 August 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.