Data-Efficient Machine Learning for Small Tabular Datasets: A Comparison Study
Authors/Creators
Description
This study compares three machine learning algorithms (Logistic Regression, Random Forest, XGBoost) across four healthcare and finance datasets with sample sizes ranging from 303 to 30,000. Evaluation metrics include accuracy, F1-score, and ROC-AUC. Key findings: Logistic Regression outperforms ensemble methods on small datasets (<500 samples), while Random Forest dominates on larger datasets. XGBoost failed to achieve best performance on any dataset. Results provide practical guidance for algorithm selection in resource-constrained environments.
Files
paper.pdf
Files
(97.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:6aadc344e0dea7499efdd8ff5a553b43
|
97.3 kB | Preview Download |