Published September 25, 2020 | Version 1.0.0

The S&M-HSTPM2d5 dataset: High Spatial-Temporal Resolution PM 2.5 Measures in Multiple Cities Sensed by Static & Mobile Devices

  • 1. Carnegie Mellon University
  • 2. Tsinghua University
  • 3. Stanford University

Description

This S&M-HSTPM2d5 dataset contains the high spatial and temporal resolution of the particulates (PM2.5) measures with the corresponding timestamp and GPS location of mobile and static devices in the three Chinese cities: Foshan, Cangzhou, and Tianjin. Different numbers of static and mobile devices were set up in each city. The sampling rate was set up as one minute in Cangzhou, and three seconds in Foshan and Tianjin. For the specific detail of the setup, please refer to the Device_Setup_Description.txt file in this repository and the data descriptor paper.

After the data collection process, the data cleaning process was performed to remove and adjust the abnormal and drifting data. The script of the data cleaning algorithm is provided in this repository. The data cleaning algorithm only adjusts or removes individual data points. The removal of the entire device's data was done after the data cleaning algorithm with empirical judgment and graphic visualization. For specific detail of the data cleaning process, please refer to the script (Data_cleaning_algorithm.ipynb) in this repository and the data descriptor paper.

The dataset in this repository is the processed version. The raw dataset and removed devices are not included in this repository.

The data is stored as a CSV file. Each CSV file which is named by the device ID represents the data that was collected by the corresponding device. Each CSV file has three types of data: timestamp as the China Standard Time (GMT+8), geographic location as latitude and longitude, and PM2.5 concentration with the unit of microgram per cubic meter. The CSV files are stored in either Static or Mobile folder which represents the devices' type. The Static and Mobile folder are stored in the corresponding city's folder.

To access the dataset, any programming language that can access CSV files is appropriate. Users can also open the CSV file directly. The get_dataset.ipynb file in this repository also provides an option of accessing the dataset. To successfully execute ipynb file, Jupyter Notebook with Python 3.0 is required. The following python library is also required:

get_dataset.ipynb:
    1. os library
    2. pandas library

Data_cleaning_algorithm.ipynb:
    1. os library
    2. pandas library
    3. datetime library
    4. math library

The instruction of installing the libraries above can be found online. After installing the Jupyter Notebook with Python 3.0 and the required libraries, users can try to open the ipynb file with Jupyter Notebook and follow the instruction inside the file. 

For questions or suggestions please e-mail Xinlei Chen <xinlei.chen@sv.cmu.edu>

Files

Cangzhou.zip

Files (138.7 MB)

Name Size Download all
md5:880ff5de7ce64c78ee198716ca079c1c
1.7 MB Preview Download
md5:da47057d7b50ec699ae7e5637efa1e7a
24.6 kB Preview Download
md5:87deabbcdf7edef906a74f27661095da
16.2 kB Preview Download
md5:1ec89018d05fdd59d2315c4853373cc2
80.0 MB Preview Download
md5:6bb5d1e3f22fc05443438c586b929dec
2.9 kB Preview Download
md5:ff7960d3ce7022455827d92710749129
6.0 kB Preview Download
md5:b7477c9cc248f9155ffdc85a189562af
56.9 MB Preview Download