Executive Guide
Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
- Author
- Aziz Shuaib Ausi
- Published
- August 13, 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
A recent study indicates a significant divergence in web crawling access policies between reputable news websites and misinformation sites, specifically concerning AI crawlers. Reputable sites are substantially more likely to use `robots.txt` files to restrict AI crawler access, with 60.0% disallowing at least one AI crawler, compared to only 9.1% of misinformation sites. This suggests a potential difference in strategic approaches to data access and content control on the web.
A recent study indicates a significant divergence in web crawling access policies between reputable news websites and misinformation sites, specifically concerning AI crawlers. Reputable sites are substantially more likely to use robots.txt files to restrict AI crawler access, with 60.0% disallowing at least one AI crawler, compared to only 9.1% of misinformation sites. This suggests a potential difference in strategic approaches to data access and content control on the web.
Why it matters
This finding highlights a critical difference in how various online entities manage access to their content, impacting the data sources used by large language models. Organizations relying on web crawling for information or developing AI models must consider these differential access policies, which may skew the datasets available for training and real-time information retrieval. This asymmetry could influence the accuracy, bias, and comprehensiveness of AI-generated responses, particularly regarding sensitive topics or information verification.
Key insights
- Reputable news websites are six times more likely than misinformation sites to disallow AI crawlers via
robots.txtfiles. - 60.0% of reputable sites restrict at least one AI crawler, while only 9.1% of misinformation sites do so.
- Reputable sites, on average, prohibit 15.5 AI user agents, whereas misinformation sites restrict fewer than one.
- The study specifically investigated whether these site types differ in their
robots.txtconfigurations related to AI crawlers.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2510.10315
Related publications
Previous
Methodologies for Improving the Quality of AI Tutoring in K-12 Education
Next
Small Data Explainer -- The impact of small data methods in everyday life
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
Executive Guide
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
Executive Guide
Ethics Practices in AI Development: An Empirical Study Across Roles and Regions
Executive Guide
Governing Agentic AI in FinTech
Executive Guide
Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
Executive Guide
Small Data Explainer -- The impact of small data methods in everyday life
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00245
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00245
- Version
- v1.0 · r0
- Issued
- 8/13/2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.