Published August 22, 2025 | Version v1

Development of a Modular Automated Software System for Data Trust Centers

Description

Data trust centers facilitate the legally secure, confidential exchange of sensitive data between data providers and data users. They implement strict protocols and governance structures to safeguard privacy and compliance, making them essential in sectors like healthcare, finance, and research. However, most data trust centers still rely on manual processes involving complex organizational, legal, and technical aspects, which can hinder efficiency and scalability. Automation of these processes offers numerous advantages, including increased speed and efficiency of workflows through standardization, a lower need for specialists, and enhanced scalability to handle growing data volumes. A critical challenge for data trust centers lies in ensuring the nonidentifiability of data. While some anonymization methods, such as simple anonymization of identifying attributes, are used, there are no widely accepted standards or best practices for ensuring that data remains nonidentifiable in a robust and reliable way. Ensuring the nonidentifiability of data currently requires time-consuming manual work. Furthermore, available software tools for setting thresholds for nonidentifiability factors like k-Anonymity, l-Diversity, and t-Closeness are often difficult to integrate into existing workflows of data trust centers. This increases the operational workload of these centers, slowing down data exchange and increasing the risk of non-compliance. To address these challenges, our goal is to develop a modular, web-based, open-source software system that minimizes manual work in data trust centers by automating time-consuming processes, such as ensuring nonidentifiability, consent management, removal of duplicate records through record linkage, and the management of data access permissions. Users will be able to activate desired modules according to their needs. The system will be designed to seamlessly integrate with existing infrastructures and workflows in data trust centers, improving operational efficiency and compliance while reducing the workload of specialized personnel. In this endeavor, we are building on existing software components such as gICS, gPAS, and E-PIX, which provide foundational capabilities for data privacy and governance. As part of our development process, we conducted a thorough requirement analysis through a questionnaire completed by 14 data trust centers, ensuring that our software system aligns with the diverse needs of these institutions. Moreover, we aim to establish best practices and standards for ensuring the nonidentifiability of personal data, based on insights from a comprehensive literature review. The best practices will support data trust centers in adopting effective techniques for ensuring nonidentifiability. Beyond improving the efficiency of data exchange workflows, the open source nature of our development process aims to promote transparency and trust within data trust centers. By providing quality standards and facilitating the automation of key processes, this system will enable faster integration of new processes and technologies, allowing data trust centers to better adapt to the evolving landscape of data privacy and security. Ultimately, our system aims to foster greater trust among stakeholders by ensuring confidentiality, integrity, and nonidentifiability of shared data, thus strengthening the role of data trust centers as trusted intermediaries in the digital economy.

Files

2025-08-26_CoRDI_MAST_CarinaBecker.pdf

Files (581.2 kB)

Name Size Download all
md5:763755f13d7d4992701f7c35d5409c1f
581.2 kB Preview Download

Additional details

Related works

Is described by
10.5281/zenodo.16920804 (DOI)

Dates

Available
2025-08-22