Published 2025 | Version v3

A Distributed Big Data Analytics Model for Traffic Crash Severity Analysis Using PySpark: A Chicago Case Study

Authors/Creators

Description

Traffic crashes impose severe human and economic costs in large metropolitan areas. Chicago, with approximately 2.7 million residents and over 100,000 vehicle crashes annually, presents an ideal case study for scalable crash analytics. This paper introduces a containerized Apache Spark framework that processes roughly 2 GB of heterogeneous crash records, including crash incidents, involved persons, vehicle details, and Vision Zero fatality data from the Chicago Police Department, within a reproducible Docker Compose environment consisting of one master node and two worker nodes. A PySpark analytical pipeline performs schema exploration, multi-table integration, temporal feature extraction, data cleaning, and geospatial feature engineering. The resulting visualizations, such as crash density maps, pedestrian and cyclist fatality distributions, age-stratified fatal crash overlays, road-surface condition layers, and hit-and-run hotspot maps, identify statistically significant spatial clusters and driver-risk profiles. The distributed cluster achieves speedup factors of 2.7 to 3.1 times over single-machine baselines, with nearly balanced workload partitioning across executor nodes. Findings reveal that fatal crashes concentrate in the near-west arterial corridor, that dry-surface conditions predominate even in fatal outcomes, and that pedestrian fatalities are disproportionately distributed in lower-income south and west-side neighborhoods. This work demonstrates that containerized Spark deployments provide a scalable, reproducible path to municipal traffic safety analytics capable of directly informing Vision Zero policy interventions.

Files

Distributed-Big Data-Analytics-Model.pdf

Files (238.6 kB)

Name Size Download all
md5:9547b9fb55e31cf221fff6d6cf5f5156
238.6 kB Preview Download