Published October 29, 2025 | Version v1

EXPLANA: A user-friendly workflow for EXPLoratory ANAlysis and feature selection in cross-sectional and longitudinal microbiome studies

Description

Challenge:

Researchers often want to identify the most important features that relate to outcome values in scientific studies. However, they face many analytic challenges, including complex longitudinal study designs, large volumes of data and literature, non-linear relationships, non-normal data distributions, and mixed data types (categorical and numerical). Furthermore, it is cumbersome to keep track of the specific analytics used to obtain a set of interesting features.

Solution:

EXPLANA (EXPLoratory ANAlysis) streamlines the feature selection process by automating many of these challenges. The necessary analytics are tracked and graphical and textual results are generated as part of a standalone interactive report.

EXPLANA is designed to identify features dependent on different contexts of change (for both observational and interventional longitudinal studies), including novel order-dependent categorical features (which capture changes that occur in a specific sequence, e.g., A_B vs. B_A).

While designed for longitudinal microbiome studies, EXPLANA also works for cross-sectional studies and is broadly useful as a general feature selection tool. Exploratory analysis should be used at different stages of discovery, especially early, to identify potential confounders, find artifacts, and elucidate important factors and complex interactions.

Files

explana-main.zip

Files (84.0 MB)

Name Size Download all
md5:1c0c233d7c012dc25049ffac4978c137
84.0 MB Preview Download

Additional details

Software

Repository URL
https://github.com/JTFouquier/explana
Programming language
Python , R
Development Status
Active

References

  • Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in neural information processing systems, 30.
  • Kursa, M. B., Jankowski, A., & Rudnicki, W. R. (2010). Boruta–a system for feature selection. Fundamenta informaticae, 101(4), 271-285.
  • Breiman, L. (2001). Random forests. Machine learning, 45(1), 5-32.
  • Hajjem, A., Bellavance, F., & Larocque, D. (2014). Mixed-effects random forest for clustered data. Journal of Statistical Computation and Simulation, 84(6), 1313-1328.