AutoML Feature Engineering for Student Modeling Yields High Accuracy, but Limited Interpretability

doi:10.5281/zenodo.5275315

Published August 26, 2021 | Version 1.0.0

Journal article Open

AutoML Feature Engineering for Student Modeling Yields High Accuracy, but Limited Interpretability

Bosch, Nigel¹

1. University of Illinois Urbana-Champaign

Automatic machine learning (AutoML) methods automate the time-consuming, feature-engineering process so that researchers produce accurate student models more quickly and easily. In this paper, we compare two AutoML feature engineering methods in the context of the National Assessment of Educational Progress (NAEP) data mining competition. The methods we compare, Featuretools and TSFRESH (Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests), have rarely been applied in the context of student interaction log data. Thus, we address research questions regarding the accuracy of models built with AutoML features, how AutoML feature types compare to each other and to expert-engineered features, and how interpretable the features are. Additionally, we developed a novel feature selection method that addresses problems applying AutoML feature engineering in this context, where there were many heterogeneous features (over 4,000) and relatively few students. Our entry to the NAEP competition placed 3rd overall on the final held-out dataset and 1st on the public leaderboard, with a final Cohen's kappa = .212 and area under the receiver operating characteristic curve (AUC) = .665 when predicting whether students would manage their time effectively on a math assessment. We found that TSFRESH features were significantly more effective than either Featuretools features or expert-engineered features in this context; however, they were also among the most difficult features to interpret based on a survey of six experts' judgments. Finally, we discuss the tradeoffs between effort and interpretability that arise in AutoML-based student modeling.

Files

-1665458234.pdf

Files (533.2 kB)

Name	Size	Download all
-1665458234.pdf md5:106d48e04d1e8b2b878c5945a1feb43a	533.2 kB	Preview Download

Additional details

Is cited by: https://jedm.educationaldatamining.org/index.php/JEDM/article/view/501 (URL)

	All versions	This version
Views	161	159
Downloads	105	104
Data volume	58.1 MB	57.6 MB

AutoML Feature Engineering for Student Modeling Yields High Accuracy, but Limited Interpretability

Creators

Description

Files

-1665458234.pdf

Files (533.2 kB)

Additional details

Related works