Knowledge Resource
The Enforcement and Feasibility of Hate Speech Moderation
- Author
- Aziz Shuaib Ausi
- Published
- 7 September 2026
- Reading time
- 1 min
- Publication type
- Knowledge Resource
- Availability
- Open access
A recent audit of hate speech moderation on Twitter (now X) found that approximately 80% of identified hateful tweets, including violent content, remained online five months after posting. The platform's removal rates for hate speech were only marginally higher than for non-hateful content, significantly lagging behind moderation of issues like scams or adult content. Current automated detection methods are insufficient for reliable classification but can effectively prioritize content for human review, suggesting that current staffing levels are a major impediment to effective enforcement.
Why it matters
The persistent and widespread presence of unmoderated hate speech on online platforms presents significant reputational, regulatory, and societal risks. Ineffective moderation undermines platform integrity, can deter users, and may invite stricter external oversight or legal action, impacting market position and public trust.
Key insights
- 80% of hateful tweets, including violent content, remained online on Twitter (now X) five months after initial posting.
- Removal rates for hate speech were only marginally higher than for non-hateful tweets.
- Moderation of hate speech was significantly less effective compared to moderation of scams or adult content.
- The severity and reach of hateful content did not significantly impact its likelihood of removal.
- Automated detection systems are currently unable to reliably classify hate speech but are effective at ranking content for human triage.
- Simulation suggests that current staffing levels for human review are insufficient to address the volume of hateful content.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2604.12289
Related intelligence and resources
Previous
TUX: Measuring Human--AI Tacit Understanding
Next
Differentiable Electricity-Market Clearing for Gradient-Based Planning
Differentiable Electricity-Market Clearing for Gradient-Based Planning
Knowledge Resource
TUX: Measuring Human--AI Tacit Understanding
Knowledge Resource
Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Unequal Protection from Perceived Toxicity
Knowledge Resource
HarmReduction: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
Knowledge Resource
AI agents reshape consensus formation in human groups
Knowledge Resource
The PIONEER Project: A PrIvacy companion for mOtivatioN and knowlEdge transfER
Knowledge Resource
Citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). The Enforcement and Feasibility of Hate Speech Moderation. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00215
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXE-2026-00215
- Version
- v1.0 · r0
- Issued
- 7 September 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.