Published December 4, 2025 | Version 1

MLDR: A Machine Learning-driven Radio Interface

Description

1. Introduction

This document contains the data management plan of the MLDR: A Machine Learning-driven Radio (MLDR) Interface project. It reports the data types the project will be dealing with, how they will be stored and released / disseminated. Special attention is placed on the open access publication policy the project will follow. The structure of this report is based on the Data Management Plan (DMP) submitted by a project partner (AGH University) to its national funding agency (i.e., National Science Centre, NCN).

MLDR is a project from the CHIST-ERA 2022 Call — Wireless Artificial Intelligence topic. It started in February 2024, and will finish in January 2027. More details about the project can be found in the project web page: https://www.upf.edu/web/mldr.eu

As indicated earlier, Publications are the most discussed and understood data. MLDR adheres to Open Access publication as a matter of principle, and compliance with Horizon Europe rules for Open Access to scientific publications needs to be ensured. For each publication we will choose the most suitable approach (either “green” OA or “gold” OA) for each publication concerned. 

The e-repository, which is the institutional repository of UPF, the coordinator of the MLDR project, is among the top 20 providers of publications and projects in H2020 and ERC: it meets all the requirements established by the European Union within the framework of Open Access publishing, and will be used for UPF co-authored publications. 

Regarding Datasets. The project will follow the principle of "as open as possible, as closed as necessary" for its data, metadata, and software under the FAIR principles to make data and software findable, accessible, interoperable and reusable. The project open metadata will be licensed under CC or equivalent. 

This DMP is released by April 2024, and will be updated—if required—by the end of the project.

2. Data Management

To make research data findable, accessible, interoperable and re-usable (FAIR), the DMP includes information on: 
a) what data will be collected, processed and/or generated; 
b) the handling of any kind of data accessible by/from the project during and after the end of the project; 
c) which methodology and standards will be applied; 
d) whether any resulting data will be shared/made open access; and 
e) how data will be curated and preserved (including after the end of the project). 

2.1 Data description and collection or re-use of existing data 

What data (for example the types, formats, and volumes) will be collected or produced? 

MLDR will generate the following types of data:
Temporal data describing the evolution of a certain protocol / functionality / system;
Use-case descriptions and KPIs; Specific data examples include: channel occupancy, packet error rate, received signal, CSI angles, etc. 
ML algorithms.

The MLDR evaluation framework (codes and models) will be published as software, including details to reproduce the same experiments done during the project execution. 

A detailed implementation of the use-cases, and datasets obtained will be also made publicly available. Datasets will be published describing how they were obtained (experimental setup and environmental conditions). In case they were obtained using the MLDR evaluation framework, we will explicitly provide the necessary scripts to generate them. An example dataset made available by the consortium is the WACA FCB dataset containing spectrum measurements.

Regarding ML algorithms, they will be published as software modules and made available to other researchers so that they can reproduce the results of the project as well as apply​ ​them​ to other problems.

The data is classified into different types, namely, publications, data sets and tools. We will also have all the documentation generated by the project. In detail, the project will create new data in the form of: 

Scientific publications. 
Tools and libraries, i.e., simulation modules, implementation of algorithms, data processing routines, etc. 
Source code, from abovementioned tools and libraries. 
Experimental (raw) results - the output of running experiments, e.g., as a CSV file with numerical values. 
Measurements from real systems, e.g., RSSI values.
Analyzed results - the output of using data analysis tools (e.g., custom Python scripts) to analyze the raw, experimental results.
Project documentation: slides and reports.
 
Any newly created data will be subject to validity tests in which the outcome is expected. For example, a newly developed simulation scenario (source code) will be run in settings where both the raw and analyzed results can be compared with results from the literature. All experiments will be run multiple times (if it proceeds) and the results will be subject to a statistical analysis to ensure their validity. 

MLDR does not require data from external sources. However, we are open to using external datasets if they can contribute to achieve the goals of the project, e.g., to pre-train a certain AI/ML module. The project will reuse publicly available open-source applications (e.g., existing simulators), libraries (e.g., matplotlib), and implementations (e.g., of machine learning algorithms). 

How data will be stored? Which data formats will be used?

Data will be stored on machines—either physical or virtual—dedicated to support the project. When required, the machine will be running a server of a distributed version control system (e.g., git) which will encompass all the data produced in the project. This system will ensure that revisions of all files are kept and that the contribution of each project member can be documented. 

Regarding the storage and curation of the research data generated and/or collected during the project, we plan the following:
Repositories: Github, Zenodo, IEEE Dataport, arXiv, TechRxiv etc.
Preservation: All listed repositories guarantee data preservation. 

For project documentation, and sharing data between partners, Google Drive will be used. It provides all access and security requirements, as well as data management services (back-up, multiple copies, version control, etc.)

We will put a strong emphasis in the project on using text-based formats, which ensure full readability: 
source code: .cc, .py 
results: .csv 
figures: .svg 
articles: .tex, .bib

These formats do not require much space, usually < 1 GB. More space is required by the open-source applications, e.g., ns-3 requires about 2 GB. Nonetheless, the project will not consume much storage capacity. 

2.2 Documentation and data quality

What metadata and documentation (for example methodology or data collection and way of organising data) will accompany data? 

The project will require storing meta-data about the experimental results, which will provide the following information: 
version of code used to run experiment, 
experiment settings. 

We will store this information using text-based files and (if possible) using dedicated open-source software, e.g., sacred, to keep track of all the information required to reproduce the results. Software such as sacred uses a dedicated database for storing results, with links to the results. For the text-based approach, we will use the following folder structure: %data-%id-%name-%config\ where the variables are, respectively, date of experiment, unique experiment id (consecutive number), experiment name, unique configuration name. Each folder will contain a human-readable README file. Furthermore, we plan to adopt an open source and text-based meta-data description specification such as Data Packages.

What data quality control measures will be used? 

The tools used to generate data in the project are mostly based on pseudo-random number generators (which means they are fully reproducible) and on deterministic data analysis. Therefore, the acquiring, processing, and analysis of data does not impact the quality of the data. Also, the obtained data does not require cleaning. 

Nonetheless, the data can be erroneous due to errors in implementation (of simulation scenarios or ML algorithms). Software testing methods and cross-validation with other tools (where available) will be used to eliminate such errors. 

2.3 Storage and backup during the research process 

How will data and metadata be stored and backed up during the research process? 

Google Drive will be used to store and backup data.

When applicable, data and meta-data will be stored on dedicated machines with a version control system. The machine will be regularly and automatically backed up using procedures provided by the hypervisor. It will allow rescuing data in case of an incident. 

How will data security and protection of sensitive data be taken care of during the research? 

Google Drive provides the required access control and security mechanisms to protect the generated data.

In case the project's machines are used, they will only be accessible through secured protocols.

In any of the two cases, access will be limited only to project personnel. 

The project does not encompass any sensitive data. 

2.4 Legal requirements, codes of conduct 

If personal data are processed, how will compliance with legislation on personal data and on data security be ensured? 

No personal data is expected to be used.

In the eventual case the use of personal data is required, it will be handled with appropriate ethical procedures. If needed for the implementation of (specific elements of) the MLDR project, the academic partners will acquire approval from their respective ethical committees. 

How will other legal issues, such as intelectual property rights and ownership, be managed? What legislation is applicable? 

We will follow the Consortium Agreement signed by MLDR partners.

2.5. Data sharing and long-term preservation 

How and when will data be shared? Are there possible restrictions to data sharing or embargo reasons? 

Google Drive will be used to store, backup, share and preserve the data, as well as the other repositories mentioned above, e.g., Zenodo.

In case of AGH University, data will be discoverable and shared in the RODBUK Cracow Open Research Data Repository. The Principal Investigator will decide about sharing data between co-investigators and outside the research group. Moreover, data sharing could be postponed because of protect intellectual property, patent procedure or before publishing – all restrictions will be made on demand and data sharing will be limited in time. 

In case of UPF, the e-Repositori — which also supports datasets — will also be used.

How will data for preservation be selected, and where will data be preserved long-term (for example a data repository or archive)? 

Google Drive will be used to store, backup, share and preserve the data, as well as the other repositories mentioned above, e.g., Zenodo.

In particular, data directly related to a given research paper will be stored in external data repositories as per the publishers recommendation. For example, IEEE (one of the main publishers in the field), provides a dedicated and free repository for authors: IEEE DataPort. Alternatively, we will use Zenodo (also recommended by IEEE), which supports the FAIR Data principles. 

A data section will exist on MLDR website, providing a central entry point, with appropriate identifiers and sufficient high quality documentation, open-source tools, data, and metadata, including licensing terms, promoting the FAIR approach as part of an active dissemination policy. 

What methods or software tools will be needed to access and use the data? 

As we will operate on text-based data only, no special software is required to access the data. 

How will the application of a unique and persistent identifier (such us a Digital Object Identifier (DOI)) to each data set be ensured? 

Data will have a digital object identifier (DOI).  

2.6 Data management responsibilities and resources 

What resources (for example financial and time) will be dedicated to data management and ensuring the data will be FAIR (Findable, Accessible, Interoperable, Re-usable)? 

All MLDR partners have allocated the required resources for data management.   

3. Open access to publications and data

We will follow the CHIST-ERA Open Science policy to provide clarity, ensure high quality, facilitate replication by third parties, and increase the visibility and adoption of our project outcomes contributing to a high impact of research. The Open Science Coordinator (AGH University) will ensure that all MDLR dissemination activities are done in accordance with this policy. In NCN-funded projects, AGH University has experience in following the NCN Open Access Policy, which is consistent with Plan S. All publications (either the Author Accepted Manuscript or Version of Record) are available under CC-BY licences and all data used in publications is publicly available as FAIR Data with a CC0 licence. 

All project scientific publications will be freely accessible to foster free access to scientific knowledge: each partner will choose the most suitable approach (either ‘green’ OA or ‘gold’ OA). The consortium has devoted a fraction of the budget to cover ‘gold’ OA costs.

We will archive publication pre-prints on the MLDR website and on repositories listed in the Directory of Open Access Repositories (OpenDOAR) such as arXiv and TechRxiv. Similarly, all other material generated in the project (news, software libraries, etc.) will be published on the webpage of the project following the same OA policies, as well as in other institutional OA repositories from the partners, such as e-Repository (UPF) - Open Aire compliant repository - and Patio (University of Oulu). In addition all four partners have agreements with many publishers including IEEE journals for open access. 

The consortium is committed not only to open the publications, but also the datasets, code, and tools resulting from MLDR for the sake of research reproducibility. Datasets, algorithms and code published in repositories such as Zenodo and GitHub will also be accessible from the project webpage. The terms of use by third parties will be defined by the General Public License. 

4. Conclusions

This report has discussed the data which the project intends to make accessible to the wider community, indicating the strategy to manage the different types of data both internally and to make data accessible. 

Files

MLDR_A_Machine_Learningdriven_Radio_Interface.json

Files (144.7 kB)

Additional details

Funding

CHIST-ERA||CHIST-ERA
MLDR - MLDR: A Machine Learning-Driven Radio Interface ANR-23-CHR4-0005