1 min readKnowledge Resource

Knowledge Resource · Open access

Research Summary: terms.txt: A Consent and Compensation Protocol for Agentic Web Access

Original authors
Attribution requires verification
Resource prepared by
Aziz Shuaib Ausi
Resource type
Research Summary / Knowledge Resource
Published
11 September 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
Checking access…

The traditional unwritten agreement governing web access, where sites allowed crawlers in exchange for traffic from search engines, is faltering due to the rise of AI crawlers and agents. These automated clients now dominate web requests, with AI platforms fetching significant volumes of data without proportional return traffic. Existing control mechanisms like robots.txt are insufficient to express identity, purpose, terms, or pricing for machine access, and can be circumvented. A new standard, terms.txt, is proposed to address these challenges by enabling path-specific, purpose-driven machine access terms, alongside a secure, origin-enforced negotiation protocol for consent and compensation.

Why it matters

This development highlights a fundamental shift in how digital information and resources are accessed and valued online, posing significant governance and operational challenges for organizations relying on web presence. Addressing the imbalance of data extraction by automated agents without adequate compensation or clear consent is critical for maintaining the economic viability and integrity of web-based operations and content creation. Establishing clear terms for machine access can protect intellectual property and ensure equitable value exchange in the evolving digital ecosystem.

Key insights

  • The unwritten bargain of the open web, where sites allowed crawlers in exchange for return visitors, is being undermined by AI crawlers and agents.
  • Automated clients now constitute the majority of web requests, with AI training dominating classified crawling activity.
  • Major AI platforms are observed to fetch thousands of pages for each visitor they return, indicating an imbalanced exchange.
  • The current robots.txt protocol is inadequate for modern machine access, lacking capabilities for expressing identity, purpose, terms, or pricing.
  • robots.txt is also prone to circumvention, and newer alternatives are often proprietary CDN features.
  • The research proposes terms.txt, a new protocol for per-path, per-purpose machine access terms.
  • terms.txt would be complemented by an origin-enforced exchange mechanism utilizing Web Bot Auth signatures, signed intent, delegation tokens, and HTTP 402 negotiation.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2609.11152

Citation

Cite the original work (APA 7)

The original source is authoritative for this citation. Cite the source publication directly — this attribution is pending verification. Open the original source.

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXE-2026-00429
Version
v1.0 · r0
Issued
11 September 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication