AI-ForestWatch Dataset: Annual Landsat-8 Multispectral Imagery and Forest Cover Change Labels for 15 Districts of Khyber Pakhtunkhwa, Pakistan, 2014 to 2020
Authors/Creators
- 1. National University of Sciences and Technology, School of Electrical Engineering and Computer Science, Islamabad, Pakistan
- 2. National Center of Artificial Intelligence, Deep Learning Laboratory, Islamabad, Pakistan
Contributors
Hosting institution (2):
Sponsor:
Description
Overview
The AI-ForestWatch dataset supports the development and evaluation of an automated framework for forest cover estimation and forest change detection using multispectral satellite imagery. The dataset was developed for pixel level classification of forest and nonforest areas and for analysing annual forest cover changes over time.
The study focuses on 15 districts of Khyber Pakhtunkhwa, Pakistan, located within regions associated with the Billion Tree Afforestation Project. The dataset brings together Landsat-8 imagery, annual satellite image composites, digitised land cover annotations, vegetation information, and forest change analysis products.
The main purpose of the dataset is to support reproducible research in forest monitoring, semantic segmentation, land cover classification, satellite image analysis, and environmental change assessment.
Geographic Coverage
The dataset covers the following 15 districts of Khyber Pakhtunkhwa:
Hangu, Karak, Kohat, Nowshehra, Battagram, Abbottabad, Kohistan, Haripur, Tor-Ghar, Mansehra, Buner, Lower Dir, Malakand, Shangla, and Swat.
These districts were selected because they contained meaningful forest cover and represented areas where substantial afforestation activity had taken place. Districts with less than one percent forest cover in the 2015 reference maps were not included. Chitral and Upper Dir were also excluded because extensive snow cover made it difficult to produce reliable forest annotations.
Temporal Coverage
The satellite imagery and annual forest analysis cover the period from 2014 to 2020.
The year 2015 serves as the reference annotation year because land cover maps were available for that year. The digitised 2015 forest and nonforest maps were used as ground truth for model training, validation, and testing.
After training on the 2015 reference data, the semantic segmentation model was applied to annual Landsat-8 composites from 2014 through 2020. This process generated a temporal sequence of forest cover maps that could be compared to identify forest gain, forest loss, and areas with no detected change.
Satellite Imagery
The remote sensing component is based on Landsat-8 top of atmosphere imagery. Landsat-8 provides imagery with a spatial resolution of approximately 30 metres per pixel for the relevant multispectral bands and a revisit period of approximately 16 days.
For each district and each year, satellite scenes with less than 10 percent cloud cover were selected where possible. When sufficiently clear scenes were unavailable, the cloud cover threshold was relaxed to 20 percent for some districts.
A pixel wise median was calculated across the selected scenes for each spectral band. This produced a clean annual composite that reduced the effects of cloud cover, atmospheric variation, and seasonal differences. The resulting composites were intended to represent the general land cover conditions for each year.
Ground Truth Development
The reference land cover maps were originally maintained in printed form by the Forest Department of Khyber Pakhtunkhwa. The original maps contained ten land cover categories, including forests, agricultural land, alpine pasture, shrubs and bushes, and other surface classes.
For this dataset, the original classes were simplified into two main categories. Forest areas were retained as the forest class. All remaining land cover categories were combined into a single nonforest class.
Google Earth Engine was used to obtain Landsat-8 imagery. QGIS was used during the digitisation and georeferencing process. Distinct geographic features, including river bends and lakes, were selected as ground control points to align the printed land cover maps with the corresponding Landsat-8 imagery.
A thin plate spline transformation was then used to map the reference data onto the satellite images. District boundary shapefiles were applied so that pixels outside each district were marked as invalid or null. The resulting annotations distinguish forest pixels, nonforest pixels, and invalid pixels outside the district boundaries.
Spectral Information
The original Landsat-8 imagery contains 11 spectral bands. These bands represent coastal and aerosol information, visible blue, visible green, visible red, near infrared, shortwave infrared, panchromatic, cirrus, and thermal infrared measurements.
Seven vegetation and environmental indices were calculated and added to the spectral information:
Normalized Difference Vegetation Index, Enhanced Vegetation Index, Soil Adjusted Vegetation Index, Modified Soil Adjusted Vegetation Index, Normalized Difference Moisture Index, Normalized Burn Ratio, and Normalized Burn Ratio 2.
Combining the 11 Landsat-8 bands with the seven calculated indices produced an augmented 18 channel input. This extended representation was designed to provide additional information about vegetation health, moisture conditions, soil influence, canopy density, and potential burned areas.
Dataset Preparation
The district images have an average size of approximately 4000 by 4000 pixels, although the exact dimensions differ between districts.
For model development, the district images were divided into patches measuring 256 by 256 pixels. A total of 3,375 patches were created for training, validation, and testing.
The data split reported in the paper consists of:
Training set, 2,700 patches, representing 80 percent of the data.
Validation set, 338 patches, representing 10 percent of the data.
Test set, 337 patches, representing 10 percent of the data.
During model training, smaller 128 by 128 pixel inputs were randomly cropped from the prepared patches.
Forest Cover and Change Products
The AI-ForestWatch workflow uses a UNet based semantic segmentation model to classify every valid pixel as forest or nonforest.
Annual forest cover maps can be compared at pixel level to identify changes between years. The analysis supports the calculation of forest cover percentage, forest gain percentage, forest loss percentage, and effective forest cover change percentage.
The change products distinguish areas where forest cover increased, areas where forest cover decreased, and areas where no change was detected. These products can support the assessment of afforestation programmes, forest degradation, environmental management, land cover planning, and long term ecological monitoring.
Potential Applications
The dataset can be used to train and evaluate semantic segmentation models for satellite imagery.
It can support research on annual forest cover mapping, afforestation assessment, deforestation monitoring, land cover classification, environmental change detection, vegetation analysis, and geospatial artificial intelligence.
It can also be used to compare deep learning methods with conventional classification approaches such as random forest, decision tree, and logistic regression.
Additional applications include the development of regional forest monitoring systems, evaluation of vegetation indices, testing of multispectral input configurations, and preparation of evidence for environmental policy and natural resource management.
Known Limitations
The 2015 reference annotations were produced by digitising printed land cover maps. The original maps had a lower visual resolution than the Landsat-8 imagery. Georeferencing the printed maps onto satellite imagery can therefore introduce differences in boundaries and estimated forest cover percentages.
The paper reports that no independent numerical accuracy assessment was performed on the original printed land cover maps. The digitised labels should therefore be treated as reference annotations rather than error free ground truth.
Only the year 2015 contains directly available reference annotations. Forest cover outputs for the remaining years were generated through model inference using annual Landsat-8 composites.
Cloud filtering was based primarily on a threshold of less than 10 percent. This threshold was relaxed to 20 percent in certain districts where sufficiently clear imagery was not available.
Some districts were excluded because of limited forest cover or extensive snow cover. The dataset should therefore not be interpreted as complete coverage of every district in Khyber Pakhtunkhwa.
Citation
Users of this dataset should cite both the Zenodo dataset record and the related journal article:
Annus Zulfiqar, Muhammad M. Ghaffar, Muhammad Shahzad, Christian Weis, Muhammad I. Malik, Faisal Shafait, and Norbert Wehn. AI-ForestWatch: semantic segmentation based end-to-end framework for forest estimation and change detection using multi-spectral remote sensing imagery. Journal of Applied Remote Sensing, Volume 15, Issue 2, Article 024518, 2021. DOI: 10.1117/1.JRS.15.024518.
Files
AI_ForestWatch.zip
Additional details
Dates
- Collected
-
2014/2020Data collection period covered by the dataset.
- Issued
-
2021-05-31Dataset version 1.0 released.