ai
The Enforcement and Feasibility of Hate Speech Moderation
arXiv: Computers and SocietyInternationalHigh confidence1 min
What changed
A recent audit of hate speech moderation on Twitter (now X) found that approximately 80% of identified hateful tweets, including violent content, remained online five months after posting. The platform's removal rates for hate speech were only marginally higher than for non-hateful content, significantly lagging behind moderation of issues like scams or adult content. Current automated detection methods are insufficient for reliable classification but can effectively prioritize content for human review, suggesting that current staffing levels are a major impediment to effective enforcement.
Why it matters
The persistent and widespread presence of unmoderated hate speech on online platforms presents significant reputational, regulatory, and societal risks. Ineffective moderation undermines platform integrity, can deter users, and may invite stricter external oversight or legal action, impacting market position and public trust.
What to watch
80% of hateful tweets, including violent content, remained online on Twitter (now X) five months after initial posting.
Forward consideration, not a verified fact.
Reported by arXiv: Computers and Society, International. The document itself is not reproduced here.
Read the original publication