Intelligence

ai

Large language models simulate intersectional synthetic identities with a budget of one to two dimensions

arXiv: Computers and SocietyInternationalHigh confidence1 min

What changed

Research from arXiv highlights that large language models (LLMs) used as synthetic survey respondents struggle to accurately represent intersectional populations. While real-world subgroup opinions become more distinctive with increasing intersectional identities, LLM-simulated respondents exhibit a 'collapse' where a single identity component often explains responses better than additive combinations of multiple identities. This suggests LLMs do not genuinely simulate complex, multi-dimensional identities.

Why it matters

This research is strategically important because it reveals a significant limitation in the current capability of large language models to accurately simulate complex human populations, particularly those with intersectional identities. This impacts the reliability and validity of insights derived from synthetic data generated by LLMs, potentially leading to flawed policy, product, or operational decisions if not properly understood and mitigated.

What to watch

Large language models are increasingly utilized as synthetic survey respondents to access rare intersectional populations cheaply.

Forward consideration, not a verified fact.

Reported by arXiv: Computers and Society, International. The document itself is not reproduced here.

Read the original publication