1 min readExecutive Guide

Executive Guide

Large language models simulate intersectional synthetic identities with a budget of one to two dimensions

Author
Aziz Shuaib Ausi
Published
28 August 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

Research from arXiv highlights that large language models (LLMs) used as synthetic survey respondents struggle to accurately represent intersectional populations. While real-world subgroup opinions become more distinctive with increasing intersectional identities, LLM-simulated respondents exhibit a 'collapse' where a single identity component often explains responses better than additive combinations of multiple identities. This suggests LLMs do not genuinely simulate complex, multi-dimensional identities.

Checking access…

Research from arXiv highlights that large language models (LLMs) used as synthetic survey respondents struggle to accurately represent intersectional populations. While real-world subgroup opinions become more distinctive with increasing intersectional identities, LLM-simulated respondents exhibit a 'collapse' where a single identity component often explains responses better than additive combinations of multiple identities. This suggests LLMs do not genuinely simulate complex, multi-dimensional identities.

Why it matters

This research is strategically important because it reveals a significant limitation in the current capability of large language models to accurately simulate complex human populations, particularly those with intersectional identities. This impacts the reliability and validity of insights derived from synthetic data generated by LLMs, potentially leading to flawed policy, product, or operational decisions if not properly understood and mitigated.

Key insights

  • Large language models are increasingly utilized as synthetic survey respondents to access rare intersectional populations cheaply.
  • Real-world data from Pew's American Trends Panel shows that subgroup opinion becomes 2.5 times more distinctive as identities intersect.
  • LLM-simulated respondents, across eight models and 21 million simulated response distributions, do not replicate this compositional effect.
  • For LLM-simulated respondents, a single feature explains a two-feature persona's responses better than an additive combination in 75-82% of subgroups.
  • The addition of a third identity feature to LLM-simulated personas provides almost no additional explanatory power for their responses.
  • This 'collapse' of intersectional distinctiveness in LLMs persists despite various prompting strategies.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.23005

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). Large language models simulate intersectional synthetic identities with a budget of one to two dimensions. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00602

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00602
Version
v1.0 · r0
Issued
28 August 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication