Skip to main content
zenodoopen

Models and Predictions for "The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction"

<p><strong>Models and Predictions</strong></p> <p>This dataset contains the trained XGBoost and EA-LSTM models and the models&#39; predictions for the paper <a href="https://github.com/gauchm/ealstm_regional_modeling"><em>The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction</em></a>.</p> <p>For each input sequence length (10, 30, 100, 270*, 365*) and each combination of model (XGBoost, EA-LSTM), training years (3, 6, 9), number of basins (13, 26, 53, 265, 531), and seed (111-888), there are five folders. Each corresponds to a random basin sample (for 531 basins there&#39;s only one folder, since it&#39;s all basins).<br> In each folder, there are three files:</p> <ul> <li><span class="math-tex">\(\texttt{model.pkl}\)</span> (XGBoost) or <em><span class="math-tex">\(\texttt{model_epoch30.pt}\)</span></em> (EA-LSTM), which stores the pickled trained model</li> <li><em><span class="math-tex">\(\texttt{xgboost_seedNNN.p}\)</span></em> or <em><span class="math-tex">\(\texttt{ealstm_seedNNN.p}\)</span></em>, which stores a pickled dictionary that maps each basin to the DataFrame of predicted and actual daily streamflow.</li> <li><span class="math-tex">\(\texttt{attributes.db}\)</span>, which stores static catchment attributes needed for inference.</li> </ul> <p>In addition to each folder, there is a SLURM submission script called <em><span class="math-tex">\(\texttt{&lt;foldername&gt;.sbatch}\)</span></em> that was used to create and evaluate the model in the folder.</p> <p>&nbsp;</p> <p>* sequence lengths 270 and 365 only contain data for EA-LSTM.</p>

ShareScore

28/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
0

Topics