Clinical Dataset Structure: A Universal Standard for Structuring Clinical Research Datasets
Authors/Creators
Description
This deposition contains the slides associated with our talk at ISMB/ECCB 2026, presented in the NIH/ODSS special track "Social and Technical Infrastructure to Advance Health Research".
This is a full-day track hosted by the National Institutes of Health (NIH) that will explore the platforms, standards, policies, and communities that enable modern biomedical research. The session highlights NIH-funded efforts to build interoperable research ecosystems, advance FAIR data and metadata practices, promote data and software reuse, and support secure, scalable, and reproducible science. More details are available at https://www.iscb.org/ismb2026/scientific-programme/nih.
We are presenting the Clinical Dataset Structure (CDS), a standardized framework for organizing multimodal clinical research data to be FAIR and AI-ready.
Abstract
Motivation
Clinical studies collect rich multimodal data, such as surveys, wearables, and eye images. There is currently no consensus on how to structure that data into a consistently organized dataset in line with the FAIR principles. As a result, clinical datasets are hard to reuse, and datasets from different studies are not interoperable, which limits the ability to pool data across studies and particularly undermines the development of AI models.
Methodology
To address this challenge in the NIH Bridge2AI-supported AI Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project, we developed the Clinical Dataset Structure (CDS). The CDS is a standardized framework for organizing clinical research data and metadata, developed by unifying and extending existing domain-specific standard structures (e.g., Brain Imaging Data Structure), metadata schemas (DataCite, ClinicalTrials.gov), and AI/ML-focused standards (Datasheet, Healthsheet).
Key results
The CDS recommends organizing multimodal clinical research datasets into one folder per data type, named according to a set convention, where applicable standard structure must then be followed within each folder. It also recommends including multiple metadata files at the root level in human-friendly (e.g., README, Healthsheet) and machine-friendly (e.g., dataset_description.json, study_description.json) formats that collectively capture over 80 metadata elements. The CDS specification is documented openly at cds-specification.readthedocs.io. The AI-READI dataset, which is structured following the CDS, has already been downloaded over 950 times and contributed to multiple publications, demonstrating real-world impact.
Impact
The impact of the CDS is already evident through the AI-READI dataset. We expect greater impact as adoption scales, a goal we are pursuing through community outreach and tooling development to lower implementation barriers. In this presentation, we will introduce the CDS, present its implementation in the AI-READI dataset, and explain how the community can adopt and build on it.
Files
ISMB20262280PatelTalk.pdf
Files
(2.0 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:85fedbab2975a1c171031ba09535064b
|
2.0 MB | Preview Download |
Additional details
Related works
- Describes
- Other: https://cds-specification.readthedocs.io/ (URL)
- Other: https://aireadi.org (URL)
- Is supplemented by
- Other: https://github.com/fairdataihub/CDS-ISMB-NIH-2026 (URL)
Funding
- National Institutes of Health
- Bridge2AI: Salutogenesis Data Generation Project 3OT2OD032644-01S2
Dates
- Other
-
2026-07-13Presentation date