No-code machine learning workflows and data exploration: nanoHUB's efforts on FAIR data
Authors/Creators
- 1. Purdue University
- 2. University of California
Description
Predictive models based on machine learning algorithms have been widely adopted across the quantitative sciences. Access barriers to advanced simulation and data tools and specialized hardware hinder progress in most scientific and engineering disciplines. While there have been advances in making research data findable, accessible, interoperable, and reusable (FAIR), scientific workflows (for machine learning, simulations, and experimental analysis) remain poorly documented and are often not accessible. Together with the need for advanced scientific computing tools and appropriate educational materials to train the next generation workforce, application of the FAIR principles on such workflows are crucial for the community’s efforts to ensure scientific reproducibility of results.
nanoHUB’s mission to make scientific software useful to the research community has been evidenced in many of the platform’s efforts to remove such barriers. For example, the introduction of Sim2Ls, a novel library that allows researchers to develop and disseminate end-to-end workflows with verified and documented inputs and outputs. Making use of nanoHUB’s Sim2Ls, tool developers can document their scientific workflows, and make them available for online simulations by publishing in nanoHUB. Many journals incentivize researchers to share scripts necessary for their analysis through repositories. Using nanoHUB’s tools, developers can not only achieve that, but also create containerized versions of their environments so potential reviewers and future researchers have access to a functional ready-to-run version of their work.
The contribution discussed in this presentation describes ongoing efforts in nanoHUB.org towards making scientific workflows and the data generate FAIR, as well as making data science tools available to non-experts. No-code machine learning (no-code ML) allows non-experts to model data and make predictions using advanced ML tools and algorithms though an intuitive graphical user interface available through a web-browser. Thus, users with little to no experience in code development can access predictive algorithms from their preprocessed data. Using a no-code ML approach both ensures that the ML pipeline contains all necessary steps for adequate testing and that it avoids many methodological pitfalls that can compromise reproducibility.
Additionally, the use of nanoHUB’s Sim2Ls also makes the workflow and the data generated more accessible as all inputs, outputs, and results are automatically indexed and cached in nanoHUB’s ResultsDB. The ResultsDB allow for efficient caching, storage, exploration, and visualization of simulation results. Scientific workflows stored in the ResultsDB may also include metadata and details about code execution. Importantly, all data generated can be queried using simple python commands to make data accessible to the community and, at the same time, allow for further investigation in different areas of interest.
Files
Gateways2022_paper_4705.pdf
Files
(606.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:548b50629e6022024f9fd10b5d6c6f00
|
508.9 kB | Preview Download |
|
md5:b79354ab453a70bd6d898ae87611f18a
|
97.4 kB | Preview Download |