Geology ReportsSearch

USGS · 70200470

Harnessing big data to rethink land heterogeneity in Earth system models

Abstract

The continual growth in the availability, detail, and wealth of environmental data provides an invaluable asset to improve the characterization of land heterogeneity in Earth system models – a persistent challenge in macroscale models. However, due to the nature of these data (volume and complexity) and computational constraints, these data are underused for global applications. As a proof of concept, this study explores how to effectively and efficiently harness these data in Earth system models over a 1/4° ( ∼ 25 km) grid cell in the western foothills of the Sierra Nevada in central California. First, a novel hierarchical multivariate clustering approach (HMC) is introduced that summarizes the high-dimensional environmental data space into hydrologically interconnected representative clusters (i.e., tiles). These tiles and their associated properties are then used to parameterize the sub-grid heterogeneity of the Geophysical Fluid Dynamics Laboratory (GFDL) LM4-HB land model. To assess how this clustering approach impacts the simulated water, energy, and carbon cycles, model experiments are run using a series of different tile configurations assembled using HMC. The results over the test domain show that (1) the observed similarity over the landscape makes it possible to converge on the macroscale response of the fully distributed model with around 300 sub-grid land model tiles; (2) assembling the sub-grid tile configuration from available environmental data can have a large impact on the macroscale states and fluxes of the water, energy, and carbon cycles; for example, the defined subsurface connections between the tiles lead to a dampening of macroscale extremes; (3) connecting the fine-scale grid to the model tiles via HMC enables circumvention of the classic scale discrepancies between the macroscale and field-scale estimates; this has potentially significant implications for the evaluation and application of Earth system models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nathaniel W. Chaney, Marjolein H. J. Van Huijgevoort, Elena Shevliakova, Sergey Malyshev, Paul C. D. Milly, Paul P. G. Gauthier, Benjamin N. Sulman. 2018-06-14. Harnessing big data to rethink land heterogeneity in Earth system models. https://doi.org/10.5194/hess-22-3311-2018

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Technical note: A low-cost approach to monitoring relative streamflow dynamics in small headwater streams using time lapse imagery and a deep learning model

Despite their ubiquity and importance as freshwater habitat, small headwater streams are under-monitored by existing stream gage networks. To address this gap, we describe a low-cost, non-contact, and low-effort method that enables organizations to monitor relative streamflow dynamics in small headwater streams. The method uses a camera to capture repeat images of the stream from a fixed position. A person then annotates pairs of images, in each case indicating which image has more apparent streamflow or indicating equal flow if no difference is discernible. A deep learning modeling framework called streamflow rank estimation (SRE) is then trained on the annotated image pairs and applied to rank all images from highest to lowest apparent streamflow. From this result a relative hydrograph can be derived. We found that our modeled relative hydrograph dynamics matched the observed hydrograph dynamics well for 11 cameras at 8 streamflow sites in western Massachusetts. Higher performance was observed during the annotation period (median Kendall's Tau rank correlation of 0.75, with a range of 0.6–0.83) than after it (median Kendall's Tau of 0.59, with range 0.34–0.74). We found that annotation performance was generally consistent across the 11 camera sites and 2 individual annotators and was positively correlated with streamflow variability at a site. A scaling simulation determined that model performance improvements were limited after 1000 annotation pairs. Our model's estimates of relative flow, while not equivalent to absolute flow, may still be useful for many applications, such as ecological modeling and calculating event-based hydrological statistics (e.g., the number of out-of-bank floods). We anticipate that this method will be a valuable tool to extend existing stream monitoring networks and provide new insights on dynamic headwater systems.

Massachusetts

Interrogating process deficiencies in large-scale hydrologic models with interpretable machine learning

Large-scale hydrologic models are increasingly being developed for operational use in the forecasting and planning of water resources. However, the predictive strength of such models depends on how well they resolve various functions of catchment hydrology, which are influenced by gradients in climate, topography, soils, and land use. Most assessments of hydrologic model uncertainty have been limited to traditional statistical methods. Here, we present a proof-of-concept approach that uses interpretable machine learning techniques to provide post hoc assessment of model sensitivity and process deficiency in hydrologic models. We train a random forest model to predict the Kling–Gupta efficiency (KGE) of National Water Model (NWM) and National Hydrologic Model (NHM) streamflow predictions for 4383 stream gauges in the conterminous United States. Thereafter, we explain the local and global controls that 48 catchment attributes exert on KGE prediction using interpretable Shapley values. Overall, we find that soil water content is the most impactful feature controlling successful model performance, suggesting that soil water storage is difficult for hydrologic models to resolve, particularly for arid locations. We identify nonlinear thresholds beyond which predictive performance decreases for NWM and NHM. For example, soil water content less than 210 mm, precipitation less than 900 mm yr −1 , road density greater than 5 km km −2 , and lake area percent greater than 10 % contributed to lower KGE values. These results suggest that improvements in how these influential processes are represented could result in the largest increases in NWM and NHM predictive performance. This study demonstrates the utility of interrogating process-based models using data-driven techniques, which has broad applicability and potential for improving the next generation of large-scale hydrologic models.

conterminous United States

Pluvial and potential compound flooding in a coupled coastal modeling framework: New York City during post-tropical Cyclone Ida (2021)

Many coastal urban areas are prone to extreme pluvial flooding due to limitations in stormwater system capacity, with the additional potential for flooding compounded by storm surge, tides, and waves. Understanding and simulating these processes can improve prediction and flood risk management. Here, we adapt the Coupled Ocean–Atmosphere–Wave–Sediment Transport modeling framework (COAWST) to simulate pluvial flooding from post-tropical Cyclone Ida (2021) in the Jamaica Bay watershed of New York City (NYC). We modify the model to capture the volumetric effects of rainfall and parameterize soil infiltration and a stormwater conveyance system as the drainage rate. We generate a spatially continuous flood map of Ida with a root-mean-square error (RMSE) of 20 cm when compared to high-water marks, useful for understanding Ida's impacts and subsequent mitigation planning. Results show that over 23 km 2 and 4621 buildings were flooded deeper than 0.3 m during Ida. Sensitivity analyses are used to study the broader risk from events like Ida (pluvial flooding) as well as potential compound (pluvial–coastal) flooding. Spatial shifting of the storm track within a typical 12 h forecast uncertainty reveals a worst-case scenario that increases this flooded area to 62 km 2 (5907 buildings). Shifting Ida's rainfall to coincide with high tide increases this flooded area by 1 km 2 , a relatively small change due to the lack of significant storm surge. The application of COAWST to this storm event addresses a broader goal of developing the capability to model compound pluvial–coastal flooding by simultaneously representing coastal storm processes such as rain, tide, waves, erosion, and atmosphere–wave–ocean interactions. The sensitivity analysis results underscore the need for detailed flood risk assessments, showing that Ida, already NYC's worst rain event, could have been even more devastating with slight shifts in the storm track.

New York