Published November 14, 2016 | Version v1

Streamlining Study Design and Statistical Analysis

  • 1. Department of Biomedical Informatics, Center for Clinical and Translational Science, University of Utah, Salt Lake City, Utah, USA
  • 2. College of Nursing, Department of Biomedical Informatics, Center for Clinical and Translational Science, University of Utah, Salt Lake City, Utah, USA
  • 3. Center for Clinical and Translational Science, University of Utah, Salt Lake City, Utah, USA
  • 4. Center for Clinical and Translational Science, Department of Internal Medicine, University of Utah, Salt Lake City, Utah, USA
  • 5. Department of Biomedical Informatics, Center for Clinical and Translational Science, Pharmacotherapy Outcome Research Center, University of Utah, Salt Lake City, Utah, USA
  • 6. Center for Clinical and Translational Science, Department of Population Health Science, University of Utah, Salt Lake City, Utah, USA

Description

Key factors causing irreproducibility of research include those related to inappropriate study design methodologies and statistical analysis1. In modern statistical practice irreproducibility could arise due to statistical (false discoveries, p-hacking, overuse/misuse of p-values, low power, poor experimental design) and computational (data, code & software management) issues2. These require understanding the processes and workflows practiced by an organization, and the development and use of metrics to quantify reproducibility.

Within the Foundation of Discovery - Population Health Research, Center for Clinical and Translational Science, University of Utah, we are undertaking a project to streamline the study design and statistical analysis workflows and processes. As a first step we met with key stakeholders to understand the current practices by eliciting example statistical projects, and then developed process information models for different types of statistical needs using Lucidchart. We then reviewed these with the Foundation’s leadership and Standards Committee to come up with ideal workflows and model, and defined key measurement points (such as those around study design, analysis plan, final report, requirements for quality checks, and double coding) for assessing reproducibility. As next steps we will use our finding to embed analytical and infrastructural approaches within the statisticians’ workflows. This will include data and code dissemination platforms such as Box, Bitbucket and GitHub, documentation platforms such as Confluence, and workflow tracking platforms such as Jira. These tools will simplify and automate the capture of communications as a statistician work through a project. Data-intensive process will use process-workflow management platforms such as Activiti3, Pegasus4 and Taverna6. These strategies for sharing and publishing study protocols, data, code and results across the spectrum7, active collaboration with the research team, automation of key steps, along with decision support will ensure quality of statistical methods and reproducibility of research.

References
1. Reproducibility and reliability of biomedical research’, organised by the Academy of Medical Sciences, BBSRC, MRC and Wellcome Trust in April 2015.
2. National Academies of Sciences, Engineering, and Medicine. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop. Washington, DC: The National Academies Press, 2016. doi:10.17226/21915.
3. Acivit, http://activiti.org
4. Pegasus, https://pegasus.isi.edu/
5. Apache Taverna. https://taverna.incubator.apache.org/
6. Peng, R. 2011. Reproducible research in computational science. Science 334(6060):1226-1227

Notes

This project is supported by NCATS UL1TR001067.

Files

Streamlining.pdf

Files (1.1 MB)

Name Size Download all
md5:719b875196a6df16c9a905bf89971d84
1.1 MB Preview Download