Knowledge Resource · Open access
MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts
- Author
- Aziz Shuaib Ausi
- Published
- 8 September 2026
- Reading time
- 1 min
- Publication type
- Knowledge Resource
- Availability
- Open access
Recent research introduces MultiGhostBench, a comprehensive multilingual benchmark for attributing long-form, Large Language Model (LLM)-generated text. Comprising 928 books across five LLMs, six languages, and three scripts, the benchmark addresses limitations of existing tools which are often restricted to English, controlled environments, or older models. Initial evaluations indicate that current attribution methods struggle with consistency and experience performance degradation when confronted with distribution shifts in domain, author, or language.
Why it matters
The development of robust methods for identifying LLM-generated content is crucial for maintaining trust in information, intellectual property, and content integrity across various sectors. The demonstrated weaknesses of current attribution methods under diverse conditions highlight a significant vulnerability that necessitates strategic investment in research and development to mitigate potential risks associated with unverified LLM output.
Key insights
- Existing LLM authorship attribution (AA) benchmarks are limited, primarily focusing on English, controlled settings, or outdated models, with short text considerations in multilingual studies.
- MultiGhostBench is a new multilingual benchmark containing 928 books generated by five recent LLMs across six languages and three scripts.
- The average length of texts in MultiGhostBench is approximately 59,000 words per book, significantly longer than previous benchmarks.
- The benchmark supports robust evaluation of attribution methods under domain, author, and language distribution shifts.
- Evaluation of representative AA methods shows no single method consistently outperforms others across all settings.
- Performance of AA methods generally degrades significantly when encountering distribution shifts.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2609.02379
Related resources
Previous
FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
Next
When Persona Attributes Improve Population Alignment in Large Language Models
When Persona Attributes Improve Population Alignment in Large Language Models
Knowledge Resource
FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
Knowledge Resource
Deal for schools: hiring supply teachers and agency workers
Knowledge Resource
Guidance: Establishing a new academy: free school presumption
Knowledge Resource
Statutory guidance: School organisation: local-authority-maintained schools
Knowledge Resource
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
Knowledge Resource
Citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00237
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXE-2026-00237
- Version
- v1.0 · r0
- Issued
- 8 September 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.