A Distributed Big Data Analytics Model for Traffic Crash Severity Analysis Using PySpark: A Chicago Case Study
Authors/Creators
Description
Traffic crashes impose severe human and economic costs in large metropolitan areas. Chicago, with approximately 2.7 million residents and over 100,000 vehicle crashes annually, presents an ideal case study for scalable crash analytics. This paper introduces a containerized Apache Spark framework that processes roughly 2 GB of heterogeneous crash records, including crash incidents, involved persons, vehicle details, and Vision Zero fatality data from the Chicago Police Department, within a reproducible Docker Compose environment consisting of one master node and two worker nodes. A PySpark analytical pipeline performs schema exploration, multi-table integration, temporal feature extraction, data cleaning, and geospatial feature engineering. The resulting visualizations, such as crash density maps, pedestrian and cyclist fatality distributions, age-stratified fatal crash overlays, road-surface condition layers, and hit-and-run hotspot maps, identify statistically significant spatial clusters and driver-risk profiles. The distributed cluster achieves speedup factors of 2.7 to 3.1 times over single-machine baselines, with nearly balanced workload partitioning across executor nodes. Findings reveal that fatal crashes concentrate in the near-west arterial corridor, that dry-surface conditions predominate even in fatal outcomes, and that pedestrian fatalities are disproportionately distributed in lower-income south and west-side neighborhoods. This work demonstrates that containerized Spark deployments provide a scalable, reproducible path to municipal traffic safety analytics capable of directly informing Vision Zero policy interventions.
Files
Distributed-Big Data-Analytics-Model.pdf
Files
(238.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:9547b9fb55e31cf221fff6d6cf5f5156
|
238.6 kB | Preview Download |