1 min readKnowledge Resource

Knowledge Resource · Open access

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

Author
Aziz Shuaib Ausi
Published
8 September 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
Checking access…

Recent research introduces MultiGhostBench, a comprehensive multilingual benchmark for attributing long-form, Large Language Model (LLM)-generated text. Comprising 928 books across five LLMs, six languages, and three scripts, the benchmark addresses limitations of existing tools which are often restricted to English, controlled environments, or older models. Initial evaluations indicate that current attribution methods struggle with consistency and experience performance degradation when confronted with distribution shifts in domain, author, or language.

Why it matters

The development of robust methods for identifying LLM-generated content is crucial for maintaining trust in information, intellectual property, and content integrity across various sectors. The demonstrated weaknesses of current attribution methods under diverse conditions highlight a significant vulnerability that necessitates strategic investment in research and development to mitigate potential risks associated with unverified LLM output.

Key insights

  • Existing LLM authorship attribution (AA) benchmarks are limited, primarily focusing on English, controlled settings, or outdated models, with short text considerations in multilingual studies.
  • MultiGhostBench is a new multilingual benchmark containing 928 books generated by five recent LLMs across six languages and three scripts.
  • The average length of texts in MultiGhostBench is approximately 59,000 words per book, significantly longer than previous benchmarks.
  • The benchmark supports robust evaluation of attribution methods under domain, author, and language distribution shifts.
  • Evaluation of representative AA methods shows no single method consistently outperforms others across all settings.
  • Performance of AA methods generally degrades significantly when encountering distribution shifts.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2609.02379

Citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00237

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXE-2026-00237
Version
v1.0 · r0
Issued
8 September 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication