Published 2026 | Version v1

Clinical Dataset Structure: A Universal Framework for Structuring Clinical Research Datasets

  • 1. ROR icon California Medical Innovations Institute
  • 2. Department of Ophthalmology, University of Washington
  • 3. Department of Bioengineering, University of Washington
  • 4. The Roger and Angie Karalis Johnson Retina Center
  • 5. John F. Hardesty MD Department of Ophthalmology and Visual Sciences, Washington University

Description

Each year, a rapidly growing volume of datasets is shared across the research community. In clinical studies, many data types are collected per participant, such as surveys, vital signs, eye images, and more. There is currently no consensus on how to organize multimodal data and include related information known as metadata. Standards like BIDS (Brain Imaging Data Structure) exist but only for structuring individual modalities. To address this challenge in the AI Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project, we developed the Clinical Dataset Structure (CDS). The CDS provides a simple and intuitive way to organize clinical research data and metadata in line with the FAIR Principles. It recommends organizing multimodal clinical research datasets into one folder per data type, named as per a set convention, where applicable standard structure must then be followed within each folder.

Abstract

Clinical studies collect rich multimodal data, such as surveys, wearables, and eye images. There is currently no consensus on how to structure that data into a consistently organized dataset. As a result, clinical datasets are hard to reuse, and datasets from different studies are not readily interoperable. To address this challenge in the AI Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project, we developed the Clinical Dataset Structure (CDS). The CDS is an open-source (CC-BY-4.0), standardized framework for organizing multimodal clinical research data and metadata developed by unifying existing domain-specific standard structures (e.g., Brain Imaging Data Structure), metadata schemas (DataCite, ClinicalTrials.gov), and AI/ML-focused standards (Datasheet, Healthsheet). The CDS recommends organizing multimodal clinical research datasets into one folder per data type, named as per a set convention, where applicable standard structure must then be followed within each folder. It also recommends including multiple metadata files at the root-level in human-friendly (e.g., README, Healthsheet) and machine-friendly (e.g., dataset_description.json, study_description.json) formats that collectively capture over 80 metadata elements. The CDS specification is documented in detailed documentation. We are also building tooling to lower the barrier to adoption. The AI-READI dataset, which is fully structured following the CDS, has already been downloaded over 950 times and contributed to multiple publications, demonstrating real-world impact. In this presentation, we will provide details about the development of the CDS, present its implementation in the AI-READI dataset, and explain how the community can adopt and build on it.

Files

poster.json

Files (824.5 kB)

Name Size Download all
md5:d5715e2be10ef00dad7b9033570daa96
13.2 kB Preview Download
md5:4c8f1a9f8edb51172b7820af5a53acd5
811.3 kB Preview Download

Additional details

Related works

Is described by
Other: https://posters.science/discover/31201 (URL)

Funding

National Institutes of Health
Bridge2AI: Salutogenesis Data Generation Project OT2OD032644

Dates

Submitted
2026-07-20
Submitted to Zenodo through Posters.science
Other
2026-07-14
Poster presentation date