Published July 27, 2026 | Version v2

Open Land Use Reference Dataset for Palm Oil Landscapes in Indonesia

Description

This dataset was developed under the Lacuna Fund-supported initiative Advancing Oil Palm Mapping in Indonesia with Social Forestry and Machine Learning. It provides a high-resolution, open-access land use reference dataset for supporting machine learning applications in land cover classification. The dataset includes wall-to-wall labeled polygons across 6×6 km grid cells, corresponding monthly satellite imagery mosaics, and a verified validation dataset derived using Collect Earth Online (CEO). The data targets key oil palm production landscapes in Riau and West Sulawesi and supports research on forest change, social forestry, and sustainable land management.

Technical info

Record 1 — Geospatial Labels (Oil Palm Landscapes, Indonesia)

Wall-to-wall, human-annotated land-use / land-cover reference polygons for oil-palm landscapes across ten provinces of Indonesia (Sumatra, Kalimantan, Sulawesi), digitised on a systematic 6 × 6 km grid. This record contains the author-created label data only. The satellite imagery the labels were drawn on is distributed separately (see Related records below), because it carries different, source-specific licenses.

License

All contents of this record are released under Creative Commons Attribution 4.0 International (CC-BY-4.0) — see LICENSE.txt. You are free to share and adapt the data for any purpose, including commercially, provided you give appropriate credit.

Cite as: Wafiq, M. W., et al. Open Land Use Reference Dataset for Palm Oil Landscapes in Indonesia. Zenodo. https://doi.org/10.5281/zenodo.15618532

Contents

record1_geospatial_labels/
├── README.md
├── LICENSE.txt                 CC-BY-4.0 legal code
├── data_dictionary.csv         field-by-field description of every label attribute
├── class_schema.csv            canonical class_ID → name, variants, oil-palm flag
├── grid_index.csv              per-grid province, island, CRS, polygon count, label year
├── grid_province_lookup.csv    grid → province → island
├── province_summary.csv        per-province polygon / area / cell totals
├── labels/
│   └── grid_XX/                (XX = 01 … 72)
│       ├── user_land_use_grid_XX.{shp,gpkg,geojson} (+ shapefile sidecars)  annotation polygons
│       ├── user_boundary_grid_XX.shp (+ sidecars)                          cell boundary + QA
│       └── land_cover_classification_grid_XX.tif (+ sidecars)              0.3 m rasterised map
└── notebooks/                  runnable benchmark notebooks + their own README
    ├── README.md               how to run the benchmark end-to-end
    ├── 00_Prepare_Shared_Tiles │ 01_AlphaEarth_Benchmark │ 02_Clay_Sentinel2_Benchmark
    ├── rf_baseline.py          classical Random-Forest baseline + shared metric code
    └── environment.yml         conda environment

Key files

  • user_land_use_grid_XX — the core deliverable: wall-to-wall polygons, provided in three mutually consistent formats (Shapefile, GeoPackage, GeoJSON). Attributes:
    • id, plotid — internal and Collect Earth Online identifiers;
    • class_ID, class_ENG, class_BAH — canonical class code and English / Bahasa names (class_ID is authoritative; class_ENG names vary across grids and are reconciled in class_schema.csv);
    • class_l1, class_l2, class_l3 — the hierarchical typology (broad category → thematic subdivision → detailed class), derived from class_ID;
    • img_date — reference-image date for the cell, propagated to its polygons;
    • source_sensor — interpretation imagery stack (campaign-level);
    • interpreter_id — annotator identifier (recorded value where present, else the team-level MoHE_intern_team).

    See data_dictionary.csv for full definitions and allowed values.

  • class_schema.csv — 17 classes. Oil palm = class_ID 1 (Palm, initial planting) and 2 (Palm, mature/young). Coconut (ID 5) is a separate class and is not oil palm.
  • land_cover_classification_grid_XX.tif — a 0.3 m rasterised land-cover product derived from the polygons. Provided for completeness; the paper's benchmark rasterises the vectors directly rather than using this file.

Coordinate reference system

The annotation polygons (user_land_use_*) are distributed in WGS 84 (EPSG:4326) in all three formats. grid_index.csv records each grid's native imagery UTM CRS (province-dependent, e.g. EPSG:32750 for West Sulawesi, EPSG:32647/32648 for Riau), which is the CRS of the user_boundary_* files and of the separately distributed imagery.

Shapefile note. The .dbf format limits field names to 10 characters, so in the Shapefile only source_sensor appears as src_sensor and interpreter_id as interp_id. The GeoPackage and GeoJSON carry the full field names. All other names are identical across formats.

Benchmark notebooks (notebooks/)

This record also ships the runnable benchmark code — the same notebooks used for the Technical Validation in the data descriptor, written as a standalone tutorial for readers with no remote-sensing or deep-learning background. See notebooks/README.md for the full walkthrough (setup, hardware, and how to run it end-to-end).

  • 00_Prepare_Shared_Tiles — downloads this labels record plus the companion imagery record and builds the shared 128×128 tiles and train/val/test split every model reuses.
  • 01_AlphaEarth_Benchmark — Google Satellite Embedding V1 (AlphaEarth) features under a U-Net head and a Random Forest.
  • 02_Clay_Sentinel2_Benchmark — Clay v1 (frozen ViT + FPN head) on Sentinel-2, paired against the AlphaEarth row.
  • rf_baseline.py — classical Random-Forest baseline and the shared metric code every model is scored with; environment.yml pins the conda environment.

The notebooks pull the annotation polygons from this record's concept DOI (10.5281/zenodo.15618531), so they always track the latest published version.

Provenance & quality

This record contains the paper's 72 canonical grid cells (Sumatra 31, Sulawesi 27, Kalimantan 14). All 72 carry wall-to-wall land-use polygons (106–7,000+ polygons each) and Planet imagery. Imagery coverage of the other sensors is not uniform: grids 03, 18, 46 have no Sentinel-2 10 m composite (so no AlphaEarth), and grid 06 has no Landsat — see the imagery record's coverage table. Labels were produced by trained interpreters in Collect Earth Online / QGIS with multi-interpreter consensus and field validation (reported Overall Accuracy ≈ 83 %, reflecting internal thematic consistency and inter-interpreter agreement). See the accompanying data descriptor for the full methodology.

Related records

Files

class_schema.csv

Files (270.0 MB)

Name Size Download all
md5:544f35966c18aa486ac8f838610d4465
734 Bytes Preview Download
md5:7ca4709dc842a05e1b9b3f7775db4814
3.5 kB Preview Download
md5:e3ccf56797b98f1beddb8acbf80af3a7
3.3 kB Preview Download
md5:60e52d1b2b0d6755731e1caadf578b39
2.2 kB Preview Download
md5:6188b528663723798398b796f3c16dcd
269.8 MB Preview Download
md5:2ab724713fdaf49e4523c4503bfd068d
18.7 kB Preview Download
md5:a4b4bb0b1c9e9615bced7bfb13a431fb
4.0 kB Preview Download
md5:a89f963f2acd46d2a62218bb62f6e9d6
139.4 kB Preview Download
md5:a3ee051b0bb54b224f26f4caf41522d1
510 Bytes Preview Download
md5:b64daae9c070e10061bd0649c3fba2ad
6.4 kB Preview Download

Additional details

Related works

Is supplement to
Dataset: 10.5281/zenodo.21483658 (DOI)
Dataset: 10.5281/zenodo.21485143 (DOI)

Funding

Meridian Institute

Software

Development Status
Active