Abstract:
Aiming at the critical issue of suboptimal prediction accuracy in urban rainfall flood under data scarcity, it uses synthetic datasets generated by InfoWorks ICM to design a multi-scenario experiment encompassing 6 dataset lengths, 4 feature combinations and 3 rainfall distribution types. The impacts of dataset length, feature combination, and rainfall distribution on urban flood forecasting performance are systematically analyzed. Results demonstrate that extending dataset length significantly enhances model performance. The prediction deviation under light rainfall and mixed rainfall conditions will decrease by over 50% when the dataset length increases from 9, 000 to 14, 400. Feature combinations exhibit a significant regulatory effect on extreme rainfall modeling. Under data-limited conditions, the standard deviation of NRMSE for heavy rainfall prediction across different feature combinations reaches 9.1. The completeness of rainfall coverage in the training set is identified as the key to improving model generalization ability. The coefficient of determination (
R2) for long-sequence prediction will reach 0.994 under optimized feature combinations when training with mixed rainfall yields the smallest deviation. Based on these findings, dataset design guidelines for small-sample hydrological modeling are proposed.