What the study found
The study found that data augmentation can improve machine learning (ML) models for simulating river flows in data-scarce catchments, especially when training data are limited. It also found that the usefulness of these approaches depends on how much data are available.
Why the authors say this matters
The authors conclude that the findings offer practical guidance for water resource engineers and modellers on when model-specific and data-dependent data augmentation strategies may be useful for river flow modelling in data-scarce regions. The study suggests this is relevant for sustainable water resources management under a changing climate.
What the researchers tested
The researchers evaluated statistical bootstrapping and physics-based data augmentation in two data-scarce Sub-Saharan African catchments with contrasting climates. They applied these methods to a Feed Forward Neural Network (FFNN) and a Long Short-Term Memory (LSTM) model, and compared their performance with the physically based Hydrologic Engineering Center Hydrologic Modeling System (HEC-HMS).
What worked and what didn't
Comparisons of standalone ML models with HEC-HMS showed data-dependent performance: HEC-HMS performed better than ML models on very limited datasets, while ML models performed better as data availability increased. Adding data augmentation improved FFNN and LSTM performance, particularly with limited training data. Limited or comparable performance was seen when longer training datasets were used, and the augmentation approaches appeared independent of model architecture and catchment hydroclimatic conditions.
What to keep in mind
The study was limited to two data-scarce catchments in Sub-Saharan Africa. The abstract does not describe additional limitations beyond the data dependence of the results and the mixed performance of augmentation with longer training datasets.
Key points
- Bootstrapping and physics-based data augmentation improved FFNN and LSTM river-flow models when training data were limited.
- HEC-HMS outperformed the machine learning models on very limited datasets.
- Machine learning models performed better as more data became available.
- With longer training datasets, augmentation showed limited or comparable performance.
- The reported augmentation effects did not depend on model architecture or catchment hydroclimatic conditions.
Disclosure
- Research title:
- Data augmentation improved some river-flow models in scarce-data catchments
- Authors:
- James Murungi, Seith N. Mugume, Ione Loots
- Institutions:
- Makerere University, University of Pretoria, University of Pretoria
- Publication date:
- 2026-07-04
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.