1 min readExecutive Guide

Executive Guide

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web

Author
Aziz Shuaib Ausi
Published
August 13, 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

A recent study indicates a significant divergence in web crawling access policies between reputable news websites and misinformation sites, specifically concerning AI crawlers. Reputable sites are substantially more likely to use `robots.txt` files to restrict AI crawler access, with 60.0% disallowing at least one AI crawler, compared to only 9.1% of misinformation sites. This suggests a potential difference in strategic approaches to data access and content control on the web.

Checking access…

A recent study indicates a significant divergence in web crawling access policies between reputable news websites and misinformation sites, specifically concerning AI crawlers. Reputable sites are substantially more likely to use robots.txt files to restrict AI crawler access, with 60.0% disallowing at least one AI crawler, compared to only 9.1% of misinformation sites. This suggests a potential difference in strategic approaches to data access and content control on the web.

Why it matters

This finding highlights a critical difference in how various online entities manage access to their content, impacting the data sources used by large language models. Organizations relying on web crawling for information or developing AI models must consider these differential access policies, which may skew the datasets available for training and real-time information retrieval. This asymmetry could influence the accuracy, bias, and comprehensiveness of AI-generated responses, particularly regarding sensitive topics or information verification.

Key insights

  • Reputable news websites are six times more likely than misinformation sites to disallow AI crawlers via robots.txt files.
  • 60.0% of reputable sites restrict at least one AI crawler, while only 9.1% of misinformation sites do so.
  • Reputable sites, on average, prohibit 15.5 AI user agents, whereas misinformation sites restrict fewer than one.
  • The study specifically investigated whether these site types differ in their robots.txt configurations related to AI crawlers.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2510.10315

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00245

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00245
Version
v1.0 · r0
Issued
8/13/2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication