Executive Guide
MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification
- Author
- Aziz Shuaib Ausi
- Published
- 28 August 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
Research identifies significant challenges in applying machine learning models for mass-shooting risk classification due to inconsistencies across public databases. A new benchmark, MASH-Bench, demonstrates that while models perform adequately on curated sources, they exhibit severe generalization failures when applied to broader, less curated datasets, such as the Gun Violence Archive (GVA). This highlights fundamental data comparability issues affecting model reliability.
Research identifies significant challenges in applying machine learning models for mass-shooting risk classification due to inconsistencies across public databases. A new benchmark, MASH-Bench, demonstrates that while models perform adequately on curated sources, they exhibit severe generalization failures when applied to broader, less curated datasets, such as the Gun Violence Archive (GVA). This highlights fundamental data comparability issues affecting model reliability.
Why it matters
This research is crucial because it exposes fundamental data quality and interoperability issues that undermine the effectiveness of advanced analytical tools in critical risk assessment domains. Understanding these limitations is vital for resource allocation, policy development, and ensuring the reliability of intelligence derived from disparate sources.
Key insights
- Public mass-shooting databases vary significantly in coverage, feature availability, and reporting practices.
- MASH-Bench is a harmonized benchmark comprising 6,968 incidents from four U.S. databases: Kaggle, Mother Jones, Stanford MSA, and the Gun Violence Archive (GVA).
- Machine learning models (Random Forest, XGBoost, LightGBM) achieve 0.68-0.89 VeryHigh-risk recall on curated sources.
- Model generalization is poor when applied to the GVA, with recall dropping to 0.20 and precision to 0.0004.
- The research investigates the degradation through controlled feature-masking ablation to pinpoint the source of failure.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.22460
Related publications
Previous
Hybrid Panels: Toward Human-AI Collaboration in Survey Research
Next
Runtime Action Interference for AI Control of AlphaStar in StarCraft II
Embedding inter- and transdisciplinary sustainability skills and knowledge development in higher education: perspectives from an innovative new degree
Executive Guide
Critical thinking as a predictor of task functionality and artificial intelligence use among university students. A PLS-SEM approach
Executive Guide
Cognitive emotion regulation as a statistical mediator of the association between autistic traits and academic performance in university students
Executive Guide
AI self-efficacy as a predictor of satisfaction with studies: the mediating role of research motivation among Peruvian University students
Executive Guide
Generative AI and linguistic creativity in digitally multilingual higher education
Executive Guide
Digital teaching and learning strategies for enhancing self-directed learning in remote ODeL environments: evidence from Zimbabwe Open University
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00615
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00615
- Version
- v1.0 · r0
- Issued
- 28 August 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.