Preprint: Hybrid physics–ML flood model outperforms NOAA’s National Water Model in retrospective hourly U.S. simulations
A new arXiv preprint reports that a hybrid flood model combining physics-based hydrology with machine learning outperformed NOAA’s operational National Water Model in retrospective hourly simulations across the contiguous United States, with the biggest gains in predicting the timing and size of rare flood peaks. The study has not been peer reviewed, and the results were not demonstrated in live forecasting.
That comparison matters because flood conditions can build and crest within hours, making accurate national-scale hourly river forecasts important for warnings and emergency response. NOAA’s National Water Model is the country’s operational large-scale hydrologic forecasting system, so it is the benchmark that matters most if researchers are arguing a new approach could improve flood prediction.
The preprint, posted on arXiv as arXiv:2609.06794v1 and announced Sept. 6, is titled “Hourly U.S.-wide flood simulation beyond the limits of traditional and data-driven models.” Its authors — Wencong Yang, Leo Lonzarich, Yalan Song, Haoyu Ji, Ming Pan, Kathryn Lawson and Chaopeng Shen of Penn State and the Center for Western Weather and Water Extremes at Scripps Institution of Oceanography, University of California, San Diego — describe their system, called δHBV2.0MTS-MC, as a national-scale hybrid model. The authors say it covers more than 800,000 river reaches in the contiguous U.S. at a median spatial resolution of 7.2 square kilometers and evaluated performance at 2,831 gauges.
In the paper’s historical test, the authors reported a median hourly Nash-Sutcliffe efficiency, or NSE — a standard hydrology accuracy metric — of 0.683 for their model, compared with 0.461 for National Water Model version 3.0. They also reported better flood-peak timing: median timing error fell from roughly seven to eight hours for NWM v3.0 to about 4.5 to 6 hours for the new model.
The rare-flood results were a major part of the paper’s case. Using the study’s event-capture criteria, the authors said their model captured 33% more floods with return periods of at least 50 years than NWM v3.0. Against recent AI-based streamflow models, the preprint reported comparable overall hourly skill but stronger performance on rare floods, including a 34% reduction in relative peak-magnitude error for floods with return periods of at least 100 years.
The study used U.S. Geological Survey hourly streamflow observations, NOAA’s Analysis of Record for Calibration meteorological forcing data and NOAA hydrologic network data. The reported training period ran from 1991 through 2003, with validation from 2004 through 2008 and an independent test from 2009 through 2018.
Those details are important because the paper is not showing current operational forecasting performance. It is a retrospective study based on historical data, not a demonstration that the model works better in real-time public forecasting. The paper also remains a preprint rather than a peer-reviewed journal article, which means the methods and conclusions have not yet gone through formal outside review.
The authors also reported weaker results in some environments, including arid basins and snow-dominated regions, compared with more humid basins. That suggests any national-scale gains may not be evenly distributed across all U.S. river systems.
In the abstract, the authors wrote that the model is “a candidate for the next-generation National Water Model” and “sets a new operational accuracy level for national-scale flood prediction.” For now, that remains the authors’ characterization of retrospective preprint results, not evidence that NOAA has adopted, endorsed or operationally validated the system.