数据匮乏下基于LSTM的雨洪预测数据集耦合机制研究

    Research on Dataset Coupling Mechanism for LSTM-Based Rainfall Flood Prediction Under Data Scarcity

    • 摘要: 针对数据匮乏条件下城市雨洪积水预测精度欠佳的痛点,利用InfoWorks ICM生成的合成数据集,设计了涵盖6种数据集长度、4种特征组合和3类降雨分布的多情景对照试验,系统分析了数据集长度、特征组合和降雨分布对城市洪涝预报性能的影响。结果表明:延长数据集长度可显著提升模型性能,数据长度由9 000增至14 400时,小雨与混合雨量条件下预测误差降幅超50%;特征组合对极端降雨建模具有显著调节作用,在数据不足条件下,不同特征组合间大雨量预测的NRMSE标准差达9.1;训练集雨量覆盖完整性是提升模型泛化能力的关键,混合雨量训练偏差最小,在优化特征组合下长序列预测决定系数可达0.994。基于此,提出了面向小样本水文建模的数据集设计准则。

       

      Abstract: Aiming at the critical issue of suboptimal prediction accuracy in urban rainfall flood under data scarcity, it uses synthetic datasets generated by InfoWorks ICM to design a multi-scenario experiment encompassing 6 dataset lengths, 4 feature combinations and 3 rainfall distribution types. The impacts of dataset length, feature combination, and rainfall distribution on urban flood forecasting performance are systematically analyzed. Results demonstrate that extending dataset length significantly enhances model performance. The prediction deviation under light rainfall and mixed rainfall conditions will decrease by over 50% when the dataset length increases from 9, 000 to 14, 400. Feature combinations exhibit a significant regulatory effect on extreme rainfall modeling. Under data-limited conditions, the standard deviation of NRMSE for heavy rainfall prediction across different feature combinations reaches 9.1. The completeness of rainfall coverage in the training set is identified as the key to improving model generalization ability. The coefficient of determination (R2) for long-sequence prediction will reach 0.994 under optimized feature combinations when training with mixed rainfall yields the smallest deviation. Based on these findings, dataset design guidelines for small-sample hydrological modeling are proposed.

       

    /

    返回文章
    返回