Published August 19, 2026 | Version v1

RoofVIP 1.0.3.0 Fix-Tile-Scene Benchmark Dataset: 2D Roof Planar Polygons, Very High-Resolution Digital Orthophotos, Digital Elevation Model, and LiDAR Point Cloud for Building Roof Reconstruction

  • 1. EDMO icon German Aerospace Center (Oberpfaffenhofen)

Description

This dataset extends the RoofVIP Benchmark Dataset into a multimodal, fixed-scene representation for building roof analysis and cross-modal benchmarking. It is derived from geospatial products provided by the Bavarian Surveying Administration (Bayerische Vermessungsverwaltung) through Bavarian Open Geodata and combines very-high-resolution RGB imagery, elevation data, LiDAR point clouds, and manually annotated roof geometry.

The dataset corresponds to RoofVIP version 1.0.3.0 and covers approximately 18 km² of the Munich metropolitan area. Whereas the original RoofVIP benchmark is organized as object-centred image-vector pairs, version 1.0.3.0 reformats the data into fixed-size scenes of 256 × 256 pixels. At the native ground sampling distance of 0.2 m, each scene covers approximately 51.2 × 51.2 m on the ground.

A scene can contain multiple buildings, surrounding background objects, and buildings intersected by the scene boundary. Partial buildings are retained rather than removed. This scene-based organization provides a consistent spatial unit for training and evaluating image-, elevation-, point-cloud-, and multimodal learning methods.

Data Modalities

Each scene may contain the following co-registered data modalities:

  • Very High Resolution RGB imagery (VHR RGB) — derived from the Digitales Orthophoto RGB 20 cm (DOP20 RGB) provided by the Bavarian Surveying Administration.
    Original source: https://geodaten.bayern.de/opengeodata/OpenDataDetail.html?pn=dop20rgb

  • Digital Surface Model (DSM) — derived from the Digitales Oberflächenmodell 20 cm (DOM20).
    Original source: https://geodaten.bayern.de/opengeodata/OpenDataDetail.html?pn=dom20

  • Digital Terrain Model (DTM) — derived from the Digitales Geländemodell 1 m (DGM1).
    Original source: https://geodaten.bayern.de/opengeodata/OpenDataDetail.html?pn=dgm1

  • LiDAR Point Cloud — point-cloud data covering the same geographic scenes and included to support multimodal and cross-modal roof reconstruction experiments.
    Original source: https://geodaten.bayern.de/opengeodata/OpenDataDetail.html?pn=laserdaten

  • Normalized Digital Surface Model (nDSM) — a derived elevation representation calculated from the DSM and DTM, i.e.

    nDSM = DSM − DTM

    Because the nDSM is generated directly from the supplied DSM and DTM products, it does not have a separate original-data download link.

  • Vector roof annotations — manually labelled building roof structures following a Level of Detail (LoD) 2.0 representation. Each individual roof plane is represented as a polygon in a common projected coordinate reference system. Polygon vertices and boundaries can consequently be used as point, edge, polygon, or graph targets for roof reconstruction tasks.

The vector annotations were initially created in ESRI Shapefile (.shp) format and were additionally converted into NumPy-based polygon representations (.npy) for machine-learning applications.

Together, the RGB, DSM, DTM, nDSM, LiDAR point-cloud, and vector-label products provide a common spatial basis for evaluating both unimodal and multimodal building roof reconstruction approaches.

 

Dataset Processing

The dataset was generated from the original Bavarian Open Geodata products and the RoofVIP roof annotations through the following processing workflow:

  1. The geographic domain was partitioned into spatially independent regions before scene generation.

  2. The original orthophotos, elevation products, and point-cloud data were spatially matched to the RoofVIP geographic domain.

  3. The data were reformatted into fixed 256 × 256 pixel scenes.

  4. Roof-plane polygons were associated with the corresponding scenes while preserving buildings intersected by scene boundaries.

  5. DSM, DTM, RGB, and LiDAR data were spatially cropped and co-registered to the corresponding scene extent.

  6. The nDSM was derived from the DSM and DTM.

  7. Manually annotated roof-plane polygons were retained as the reference vector labels and converted into machine-learning-compatible representations where required.

  8. The resulting files were organized using consistent scene identifiers so that raster, point-cloud, derived elevation, and vector products can be directly associated across modalities.

Within each geographic partition, scenes are generated with 25% forward and side overlap, corresponding to a stride of 75% of the scene size. Scenes containing no annotated building geometry are excluded from the benchmark.

No additional geometric modification is applied to the source DSM, DTM, or LiDAR measurements beyond the spatial processing required for scene extraction, alignment, and dataset organization.

 

Training, Validation, and Test Split

For reproducible model development and evaluation, the dataset provides predefined training, validation, and test subsets using an approximately 80–10–10 split:

  • Training: 8,195 scenes

  • Validation: 1,022 scenes

  • Test: 1,028 scenes

This results in a total of 10,245 image-vector scene pairs.

The split follows the structural-complexity-based strategy used in RoofVIP. Rather than assigning scenes purely at random, descriptors representing building roof geometry and graph topology are used to characterize roof-structure complexity. These descriptors are reduced using Principal Component Analysis (PCA), producing a PCA-based representation of roof complexity that is used to balance the structural-complexity distribution across the training, validation, and test subsets.

Importantly, the geographic area is partitioned before tiling. This prevents overlapping regions, spatially adjacent duplicate content, or the same individual building from appearing in more than one subset. Scene generation is then performed independently within each partition. This procedure reduces spatial leakage while retaining comparable distributions of roof-structure complexity across the three subsets.

The predefined split should therefore be used when reporting benchmark results to facilitate reproducibility and direct comparison between methods.

 

Relationship to the Original RoofVIP Benchmark

The original RoofVIP benchmark consists of manually annotated two-dimensional roof-plane polygons and very-high-resolution RGB orthophotos organized primarily around individual building objects. Version 1.0.3.0 retains the same underlying geographic domain and roof annotations while introducing the fixed-scene representation and additional co-registered modalities.

The principal extensions are therefore:

  • conversion from object-centred samples to fixed 256 × 256 pixel scenes;

  • inclusion of DSM and DTM elevation data;

  • derivation of corresponding nDSM data;

  • inclusion of LiDAR point-cloud data for multimodal experiments;

  • preservation of manually annotated LoD 2.0 roof-plane vector geometry;

  • standardized correspondence among all modalities using common scene identifiers; and

  • provision of a predefined, geographically separated and PCA-complexity-balanced 80–10–10 train/validation/test split.

These additions enable RoofVIP to support not only RGB-based roof reconstruction but also systematic comparison of image, elevation, point-cloud, and multimodal approaches under a common benchmark configuration.

 

Licensing

The original Bavarian Open Geodata products used to construct this dataset are distributed under the Creative Commons Attribution 4.0 International License (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/

The image tiles, elevation crops, and point-cloud subsets contained in this dataset remain derived from their respective original Bavarian Open Geodata products and remain subject to the applicable CC BY 4.0 attribution requirements.

The manually produced roof annotations, converted polygon representations, derived nDSM data, scene organization, dataset curation, and associated benchmark structure are distributed as part of this dataset under the Creative Commons Attribution 4.0 International License (CC BY 4.0).

Users of the dataset must provide appropriate attribution to:

  1. the Bavarian Surveying Administration (Bayerische Vermessungsverwaltung) as the provider of the underlying geospatial data; and

  2. the authors of the RoofVIP dataset for the roof annotations, processing, dataset organization, and benchmark preparation.

Appropriate credit to the original data providers is maintained throughout the dataset. No endorsement by the original licensors is implied.

When using or redistributing the dataset, users should preserve the original source and licensing information and cite the corresponding RoofVIP publication where applicable.

Files

RoofVIP_1.0.3.0_FixTileScene_Size256.zip

Files (8.7 GB)

Name Size
md5:1f904a8c1ff0ae62a592327d318a8df6
8.7 GB Preview Download

Additional details