1 min readExecutive Guide

Executive Guide

MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification

Author
Aziz Shuaib Ausi
Published
28 August 2026
Reading time
1 min
Publication type
Executive Guide
Availability
Open access

Executive Summary

Research identifies significant challenges in applying machine learning models for mass-shooting risk classification due to inconsistencies across public databases. A new benchmark, MASH-Bench, demonstrates that while models perform adequately on curated sources, they exhibit severe generalization failures when applied to broader, less curated datasets, such as the Gun Violence Archive (GVA). This highlights fundamental data comparability issues affecting model reliability.

Checking access…

Research identifies significant challenges in applying machine learning models for mass-shooting risk classification due to inconsistencies across public databases. A new benchmark, MASH-Bench, demonstrates that while models perform adequately on curated sources, they exhibit severe generalization failures when applied to broader, less curated datasets, such as the Gun Violence Archive (GVA). This highlights fundamental data comparability issues affecting model reliability.

Why it matters

This research is crucial because it exposes fundamental data quality and interoperability issues that undermine the effectiveness of advanced analytical tools in critical risk assessment domains. Understanding these limitations is vital for resource allocation, policy development, and ensuring the reliability of intelligence derived from disparate sources.

Key insights

  • Public mass-shooting databases vary significantly in coverage, feature availability, and reporting practices.
  • MASH-Bench is a harmonized benchmark comprising 6,968 incidents from four U.S. databases: Kaggle, Mother Jones, Stanford MSA, and the Gun Violence Archive (GVA).
  • Machine learning models (Random Forest, XGBoost, LightGBM) achieve 0.68-0.89 VeryHigh-risk recall on curated sources.
  • Model generalization is poor when applied to the GVA, with recall dropping to 0.20 and precision to 0.0004.
  • The research investigates the degradation through controlled feature-masking ablation to pinpoint the source of failure.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2608.22460

Download & citation

Cite this publication (APA 7)

Aziz Shuaib Ausi (2026). MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00615

Verification

This is an authenticated institutional record.

Verification ID
ASA-EXG-2026-00615
Version
v1.0 · r0
Issued
28 August 2026
Publisher
Aziz Shuaib Ausi
Licence
All rights reserved. Reproduction requires written permission.

Verify this publication