Published November 15, 2019 | Version v3

Models and Predictions for "The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction"

  • 1. University of Waterloo

Description

Models and Predictions

This dataset contains the trained XGBoost and EA-LSTM models and the models' predictions for the paper The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction.

For each input sequence length (10, 30, 100, 270*, 365*) and each combination of model (XGBoost, EA-LSTM), training years (3, 6, 9), number of basins (13, 26, 53, 265, 531), and seed (111-888), there are five folders. Each corresponds to a random basin sample (for 531 basins there's only one folder, since it's all basins).
In each folder, there are three files:

  • \(\texttt{model.pkl}\) (XGBoost) or \(\texttt{model_epoch30.pt}\) (EA-LSTM), which stores the pickled trained model
  • \(\texttt{xgboost_seedNNN.p}\) or \(\texttt{ealstm_seedNNN.p}\), which stores a pickled dictionary that maps each basin to the DataFrame of predicted and actual daily streamflow.
  • \(\texttt{attributes.db}\), which stores static catchment attributes needed for inference.

In addition to each folder, there is a SLURM submission script called \(\texttt{<foldername>.sbatch}\) that was used to create and evaluate the model in the folder.

 

* sequence lengths 270 and 365 only contain data for EA-LSTM.

Files

README_xgboost_vs_ealstm_models_and_predictions.md

Files (18.8 GB)

Name Size
md5:3979ede59c37f47a216543d0f8ba2ef6
1.2 kB Preview Download
md5:7ec089fee84431376199762e48ba78fd
18.8 GB Download