Geology ReportsSearch

SEARCH · Geology Reports

Results for “Machine Learning with Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Applications of deep convolutional neural networks to predict length, circumference, and weight from mostly dewatered images of fish

Simple biometric data of fish aid fishery management tasks such as monitoring the structure of fish populations and regulating recreational harvest. While these data are foundational to fishery research and management, the collection of length and weight data through physical handling of the fish is challenging as it is time consuming for personnel and can be stressful for the fish. Recent advances in imaging technology and machine learning now offer alternatives for capturing biometric data. To investigate the potential of deep convolutional neural networks to predict biometric data, several regressors were trained and evaluated on data stemming from the FishL™ Recognition System and manual measurements of length, girth, and weight. The dataset consisted of 694 fish from 22 different species common to Laurentian Great Lakes. Even with such a diverse dataset and variety of presentations by the fish, the regressors proved to be robust and achieved competitive mean percent errors in the range of 5.5 to 7.6% for length and girth on an evaluation dataset. Potential applications of this work could increase the efficiency and accuracy of routine survey work by fishery professionals and provide a means for longer‐term automated collection of fish biometric data.

Illinois, Michigan, Ohio

Computational approaches improve evidence synthesis and inform broad fisheries trends

Addressing ecological impacts with effective conservation actions requires information on the links between human pressures and localized responses. Understanding links is a priority for many conservation contexts, including the world's fresh waters, which face intensifying threats to disproportionately high species diversity, including more than half of the world's fish species. Literature synthesis can uncover links and highlight potential research gaps, yet can be very cumbersome and time consuming. Emerging tools like text mining can improve efficiency in extracting relevant information from vast scientific outputs. This study synthesizes evidence of direct anthropogenic threats to major inland fisheries and examines driver-impact-response patterns using coupled automated and manual text classification methods. We screened 9336 abstracts from 45 river basins of high importance to inland fish production; 1152 abstracts contained evidence of direct threats to fish. The most common documented drivers were pollution, dams, and fishing pressure, which were most strongly linked to decreased fitness, altered reproduction, and mortality, respectively. Strong impact-response links to pollution signal potential bias toward documenting acute threats that generate more visible and immediate impacts. The use of machine learning-based text classification performed best in classifying extraneous information. Results can inform the development of inland fisheries indicators and threat-based metrics, highlight possible evidence gaps in linking global drivers to fishery-level responses, and illustrate the application of a coupled synthesis approach for improved efficiency and extraction of information relevant to conservation outcomes. The associated user and interpretation guides address accessibility and technical barriers faced by conservation scientists to improve efficiency in evidence synthesis.

Conservation Science and Practice

Leveraging artificial intelligence and machine learning to advance Chesapeake Bay research and management: A review of status, challenges, and opportunities

The Chesapeake Bay and its watershed (hereafter “Chesapeake Bay region”) have been the focus of extensive restoration efforts for several decades. These restoration efforts are guided by the Chesapeake Bay Watershed Agreement (Chesapeake Executive Council 2014) which outlines 10 goals and 31 measurable outcomes. The Chesapeake Bay is globally recognized as a model for coastal restoration due to long-term investments in monitoring, modeling, implementation and research by the Chesapeake Bay Program (CBP) partnership. These monitoring network spans tidal and non-tidal regions and provides data across multiple scales. Artificial intelligence (AI), particularly machine-learning (ML) and deep learning (DL), has emerged as a powerful tool for analyzing large, complex datasets. These techniques have gained widespread adoption across various disciplines, including ecology, hydrology, and environmental science. In the Bay context, AI/ML is increasingly being used to explore drivers of environmental change, analyze system dynamics, and predict conditions in areas with limited monitoring. The CBP partnership, particularly its Scientific and Technical Advisory Committee (STAC), has increasingly recognized the growing role of AI/ML in watershed and estuarine management. Recent Chesapeake Community Research Symposium sessions and initiatives such as the Chesapeake Global Collaboratory highlight increasing regional momentum to apply big data and AI/ML for environmental solutions. Together, these developments underscore the timely need to explore how AI/ML can help advance Chesapeake Bay restoration and management. This STAC workshop, titled “Leveraging Artificial Intelligence and Machine learning to Advance Chesapeake Bay Research and Management: A review of status, challenges, and opportunities,” was held from February 24-25, 2025, in Edgewater, Maryland to bring together over 50 federal, state, and academic scientists and partners to synthesize the current state of AI/ML applications and identify research gaps in Chesapeake Bay research and management. The workshop focused on three main objectives: 1. Summarize recent AI/ML applications and lessons learned in both tidal and nontidal areas of the Chesapeake Bay region. 2. Identify challenges and gaps in applying AI/ML approaches to Chesapeake Bay data. Such challenges and gaps may include data limitations, harmonization issues, ineffective communication of AI/ML insights, and a lack of coordination among research and management institutions. 3. Develop recommendations and identify opportunities for leveraging AI/ML to address issues across the Chesapeake Bay region. Key areas of focus may include generating new information to support watershed management, delivering AI/MLgenerated insights to managers in a clear and actionable way, and fostering greater collaboration among stakeholders within the CBP Partnership. Workshop participants engaged in science presentations and breakout sessions to develop recommendations for advancing the integration of AI/ML techniques into research and management across the Chesapeake Bay region. By synthesizing current applications, identifying challenges, and exploring new opportunities, the workshop has provided valuable insights and recommendations for better leveraging AI/ML approaches to support the success of Bay restoration efforts. Together, these recommendations provide a roadmap for enhancing data-driven, science-based decision making aligned with the goals and outcomes of the Chesapeake Bay Watershed Agreement.

Delaware, Maryland, Virginia

Physics-guided recurrent neural networks for predicting lake water temperature

This chapter presents a physics-guided recurrent neural network model (PGRNN) for predicting water temperature in lake systems. Standard machine learning (ML) methods, especially deep learning models, often require a large amount of labeled training samples, which are often not available in scientific problems due to the substantial human labor and material costs associated with data collection. ML models have found tremendous success in several commercial applications, e.g., computer vision and natural language processing. The chapter presents PGRNN as a general framework for modeling physical processes in engineering and environmental systems. The proposed PGRNN explicitly incorporates physical laws such as energy conservation or mass conservation. In particular, researchers started pursing this direction by using residual modeling, where an ML model is learned to predict the errors, or residuals, made by a physics-based model. Advanced ML models, especially deep learning models, often require a large amount of training data for tuning model parameters.

Book chapter

Assessing decadal-scale coastal change likelihood to define the accuracy and application of scientific information

Defining the accuracy and uncertainties of scientific data products is critical to the usability and trustworthiness of scientific information for environmental management and conservation purposes, such as coastal resource prioritization, design, adaptation, and mitigation. The U.S. Geological Survey has a new decadal-scale coastal change assessment product that synthesizes nearly two dozen coastal datasets. A supervised machine-learning framework is used to combine existing datasets that describe the landscape and the hazards that affect it to determine the coastal change likelihood (CCL) in the coming decade at a resolution of 10 m per pixel for the NE United States from Maine to Virginia. Here, results from a series of statistical tests conducted on source data, the supervised classification, and the CCL outcomes as compared with historical land-cover change are presented. The overall accuracy of the aggregated land-cover dataset that serves as the foundation to which other source datasets are appended is 94%. The supervised learning classification that determines the final CCL output has an overall accuracy of 92%. The CCL predictions of high expected coastal change were consistent with 95% of the coastal and low-elevation landscape change in the last 20 years, as recorded by the Coastal Change Analysis Program land-cover change atlas. Results suggest that CCL provides accurate estimates of coastal landscape change in the next decade that are consistent with recent observed change. Additionally, best practices for applying CCL for planning purposes are outlined, and citing limitations, knowledge gaps, and opportunities for improved accuracy and further investigation are considered.

Journal of Coastal Research

Forecasting water levels using machine (deep) learning to complement numerical modelling in the southern Everglades, USA

Water level is an important guide for water resource management and wetland ecosystems, defining one of the most basic processes in hydrology. This research seeks to investigate the possibility of complementing numerical modeling with a Machine Learning (ML) model to forecast daily water levels in the southern Everglades in Florida, USA. An exact analytical solution to water level may not be possible, but using the computational methods afforded by ML, the traditional numerical techniques may be enhanced to generate more robust, scalable predictions. Five locations were chosen for application of the Time-Delayed Neural Network (TDNN) and Long-Short Term Memory Recurrent Neural Network (LSTM-RNN) ML models, which were built to estimate water level with 1, 2, 3, 7 and 10 day forecasts using a simulation step of 1 day. The results showed that rainfall forecasts from weather models could improve water-level forecasts if the accuracy and performance of the weather models can be improved. The ML models presented here improve water-level predictions from a historical hydrologic model for a 24 hour forecast horizon.

Florida

Machine learning approaches to identify lithium concentration in petroleum produced waters

Prices for battery-grade lithium have increased substantially since 2020, which is propelling the search for additional sources of this important element. Battery-grade lithium is predominately recovered from continental brines. Most crude oil and natural gas wells recover briny formation water, which may represent an additional source. Chemical analysis of these waters has been shown to indicate the presence of varying concentrations of lithium and related elements. This paper briefly reviews developments and literature supporting the presence of lithium in petroleum reservoir brines. It also describes the coverage and distribution of lithium data analyses in the United States Geological Survey National Produced Waters Geochemical Database (PWGD). It then addresses the question as to whether a lithium concentration can be accurately predicted using constituents of ion chemistry in produced brines from specific geologic formations. Four machine learning algorithms are employed to classify the commercial potential of lithium in oil field brines using data from oil wells recovering formation water from the Smackover Formation. The calibrated classification models are further applied to new (out-of-sample) data from the Marcellus Formation in the Appalachian Basin. Among the approaches considered, the predictive performance and wider applicability of the gradient boosted tree and the deep neural network models are determined to be the most promising. Finally, we discuss how the calibrated models could be applied to assure the quality of the data reported from chemical laboratory analysis and for imputation when lithium values are missing.

Mineral Economics

Invertibility aware integration of static and time-series data: An application to lake temperature modeling

Accurate predictions of water temperature are the foundation for many decisions and regulations, with direct impacts on water quality, fishery yields, and power production. Building accurate broad-scale models for lake temperature prediction remains challenging in practice due to the variability in the data distribution across different lake systems monitored by static and time-series data. In this paper, to tackle the above challenges, we propose a novel machine learning based approach for integrating static and time-series data in deep recurrent models, which we call Invertibility-Aware-Long Short-Term Memory(IA-LSTM), and demonstrate its effectiveness in predicting lake temperature. Our proposed method integrates components of the Invertible Network and LSTM to better predict temperature profiles (forward modeling) and infer the static features (i.e., inverse modeling) that can eventually enhance the prediction when static variables are missing. We evaluate our method on predicting the temperature profile of 450 lakes in the Midwestern U.S. and report relative improvement of 4% to capture data heterogeneity and simultaneously outperform baseline predictions by 12% when static features are unavailable.

Conference Paper

Random forest

This entry defines and discusses the random forest machine learning algorithm. The algorithm is used to predict class or quantities for target variables using values of a set of predictor variables. It uses decision trees that are generated from bootstrap sampling of the training data set to create a "forest". The entry discusses the algorithm steps, the interpretative tools of the resulting model, current areas of research, and its limitations. Applications to the quantitative geosciences are reviewed as well as availability of software to implement the algorithm.

Book chapter

SKHASH: A python package for computing earthquake focal mechanisms

We introduce a Python package for computing focal mechanism solutions. This algorithm, which we refer to as SKHASH, is largely based on the HASH algorithm originally written in Fortran over 20 yr ago. HASH innovated the use of suites of solutions, spanning the expected errors in polarities and takeoff angles, to estimate focal mechanism uncertainty. SKHASH benefits from new features with flexible input formats and allows users to take advantage of recent advances in constraining focal mechanisms for small magnitude or poorly recorded earthquakes. The 3D locations of earthquakes and the velocity models used are varied when finding acceptable solutions. As a result, source–receiver azimuths are reflective of errors from the earthquake locations and velocity models, in addition to the takeoff angles. Users can consider weighted P ‐wave first‐motion polarities derived from traditional or machine‐learning picks, cross‐correlation consensus, and/or imputation techniques using SKHASH. Focal mechanism solutions can also be further constrained using traditional, machine learning, and/or cross‐correlation consensus S / P amplitude ratios. With improved reporting of individual and collective P polarity and S / P amplitude misfits, users can better evaluate the success of the solutions and the quality of the measurements. The reporting also makes it easier to identify potential issues with metadata, including incorrectly reported station polarity reversals. In addition, by leveraging vectorized operations, taking advantage of an efficient backend Python C Application Programming Interface, and the use of a parallel environment, the Python SKHASH routine may compute mechanisms quicker than the HASH routine.

Seismological Research Letters

Estimating traffic volume and road age in Wyoming to inform resource management planning: An application with wildlife-vehicle collisions

Road networks and their associated vehicular traffic disturb many terrestrial systems, but inventories of roads used to assess these effects often focus on the ‘where’ (e.g., local road type and density) and neglect the ‘when’ (e.g., temporal disturbance) or ‘how much’ (e.g., traffic volume disturbance). We developed annual estimates of the ‘when’ (road age) and ‘how much’ (vehicular traffic volume) across 148,172 km of highways, arterials, collectors, local, and gravel/graded roads within the state of Wyoming for the years 1986 to 2020 to provide a comprehensive dataset for future ecological investigations. We leveraged a suite of ancillary data on surface disturbances (e.g., oil & gas drilling operations, wind turbines, and open pit mines) with known establishment dates and combined them using graph theory and centrality metrics to estimate the age of each road. We then predicted traffic volume obtained from the Wyoming Department of Transportation for each year across Wyoming using a machine learning method, XGBoost, and a separate set of spatial covariates hypothesized to explain traffic patterns across large regions. We found that 132,476 km of these roads likely existed before 1986, but that 16,693 km (10.7 %) of roads have been built since 1986. Overall, our estimates of road age were 89 % accurate when assessed on a subset of 1,330 roads with high-resolution aerial imagery. Mean absolute error for predicting traffic volume ranged from 35.2 to 77.9 annual average daily traffic (aadt) for trucks and 269.2 to 516.7 aadt for all-vehicles across the 35 years. We found that mean traffic volume across the state increased by 23 % for both truck-only traffic and all vehicular traffic from 1986 to 2020. However, changes in traffic volume have varied substantially across the state (e.g., 100 % increases in volume in some areas, while other areas experienced declines of up to 1,786 %). We also illustrate a novel application of these data by predicting rates of reported wildlife-vehicle collisions (WVCs) along a subset of roads. We found evidence of a non-linear relationship that supported a threshold hypothesis for WVCs, wherein increases in traffic volume equate to increases in WVCs up to a threshold, above which increases in traffic volume result in declines in WVCs. The data provided here will enable better-informed studies of road ecology to address how roads may affect wildlife populations and key ecosystems across Wyoming.

Wyoming

Context matters: Using reinforcement learning to develop human-readable, state-dependent outbreak response policies

The number of all possible epidemics of a given infectious disease that could occur on a given landscape is large for systems of real-world complexity. Furthermore, there is no guarantee that the control actions that are optimal, on average, over all possible epidemics are also best for each possible epidemic. Reinforcement learning (RL) has been used to develop machine-readable context-dependent solutions for complex problems with many possible realisations ranging from video-games to the game of Go. RL could be a valuable tool to generate context-dependent policies for outbreak response, though translating the resulting policies into simple rules that can be read and interpreted by human decision-makers remains a challenge. Here we illustrate the application of RL to the development of context-dependent outbreak response policies to minimise outbreaks of foot-and-mouth disease. We show that control based on the resulting context-dependent policies, which adapt interventions to the specific outbreak, result in smaller outbreaks than static policies. We further illustrate two approaches for translating the complex machine-readable policies into simple heuristics that can be evaluated by human decision-makers.

Philosophical Transactions of the Royal Society B:

Global cropland-extent product at 30-m resolution (GCEP30) derived from Landsat satellite time-series data for the year 2015 using multiple machine-learning algorithms on Google Earth Engine cloud

Executive Summary Global food and water security analysis and management require precise and accurate global cropland-extent maps. Existing maps have limitations, in that they are (1) mapped using coarse-resolution remote-sensing data, resulting in the lack of precise mapping location of croplands and their accuracies; (2) derived by collecting and collating national statistical data that are often subjective, leading to substantial uncertainties in cropland-area estimates, as well as their locations; and (3) extracted from one or more classes of a land use–land cover product in which cropland classes are not the focus of mapping, leading to their mixing with other classes and creating significant errors of omission and commission. These limitations can be overcome by producing high-resolution cropland-extent maps using satellite-sensor data, such as Landsat 30-m resolution or higher. The most fundamental cropland product is the high-resolution cropland-extent map because all higher level cropland products, such as crop-watering method (that is, whether crops are irrigated or rainfed), crop types, cropping intensities, cropland fallows, crop productivity, and crop-water productivity, are dependent on a precise and accurate cropland-extent product. Given these realities, the overarching goal of this study was to produce a Landsat satellite-derived global cropland-extent product at 30-m resolution. The work, which involved a paradigm shift in how global cropland-extent maps are produced, involved the following five key steps: (1) petabyte-scale computing that involved multiyear, 8- to 16-day, time-series Landsat 30-m resolution data for the global land surface; (2) composition of analysis-ready data (ARD) cubes; (3) creation of a large global-reference data hub for machine learning; (4) use of multiple machine-learning algorithms (MLAs) by writing software and computing in the cloud; and (5) Google Earth Engine (GEE) cloud computing. The five key steps involved nine distinct phases. First, the world was segmented into 74 agroecological zones (AEZs). Second, Landsat 8- to 16-day data were used to time-composite 10-band (blue, green, red, near-infrared, short-wave infrared band 1, short-wave infrared band 2, thermal infrared, enhanced vegetation index, normalized difference water index, and normalized difference vegetation index) Landsat 30-m resolution data cubes for every 2- to 4-month time period during 3- to 4-year periods (stated as nominal-year 2015 or, simply, 2015), along with two additional 30-m resolution bands (Shuttle Radar Topography Mission elevation, and slope) in each of the 74 AEZs. Third, more than 100,000 reference-training data samples were collected using ground data (some of which were collected using a mobile application), as well as submeter- to 5-m-resolution, very high-resolution imagery sourced from other reliable sources. Fourth, reference-training data were used to create a knowledge base for separating cropland from noncropland. Fifth, MLAs such as the pixel-based supervised random forest and support-vector machines were written on the GEE using Python and JavaScript. Sixth, object-based recursive hierarchical segmentation algorithm was used, in addition to MLAs, to overcome uncertainties. Seventh, MLAs used the knowledge base to classify and separate cropland from noncropland. Eighth, accuracy assessment was conducted by generating error matrices for each of the 74 AEZs using 19,171 independent validation-data samples. Ninth, cropland areas were computed for all countries of the world and compared with United Nation’s (UN’s) Food and Agricultural Organization (FAO) and other national statistics. The outcome was a Landsat-derived global cropland-extent product at 30-m resolution (GCEP30), which has an overall accuracy of 91.7 percent. For the cropland class, producer’s accuracy was 83.4 percent, and user’s accuracy was 78.3 percent. GCEP30 calculated (using direct pixel count) the global net-cropland area (GNCA) for the year 2015 as 1.873 billion hectares (~12.6 percent of the Earth’s terrestrial area). The continental cropland distribution as a percentage of GNCA was Asia, 33 percent; Europe, 25.5 percent; Africa, 16.7 percent; North America, 14.4 percent; South America, 8.1 percent; and Australia and Oceania, 2.4 percent. The worldwide cropland areas in GCEP30 for 2015 were higher by 236 to 299 million hectares (Mha) compared to national statistics reported elsewhere for the same year (for example, in Food and Agriculture Organization’s corporate statistical database [FAOSTAT] and in the monthly irrigated and rainfed crop areas [MIRCA] database). The global cropland area reported for 2015 increased by 344 Mha (22.5 percent), compared to the year 2000. During the same period (2000–2015), the world’s population increased by 20 percent. Whereas some of these areal increases are real increases in cropland areas, others are due to the types of data, methods, and approaches used. Using the highest known resolution (compared to previous coarse-resolution global products) enabled this study to capture fragmented croplands. Coarse-resolution data compute areas on the basis of subpixels, which, for a large proportion of certain land use–land cover classes, will show only a certain percentage of the total pixel area as actual area. Subpixel areas can lead to substantial uncertainties in area computation, as determining the exact fraction of cropland areas within a coarse-resolution pixel is resource intensive and subject to errors. Other innovations in GCEP30 include reference-data hubs, machine learning, and cloud computing. Cropland areas in 214 countries, territories, departments, and regions were calculated for the year 2015 using GCEP30, on the basis of UN’s global administrative unit layers (GAUL) boundaries. The 10 leading countries in terms of cropland area (as a percentage of the GNCA) were India (9.6 percent), United States (8.95 percent), China (8.82 percent), Russia (8.32 percent), Brazil (3.42 percent), Ukraine (2.32 percent), Canada (2.29 percent), Argentina (2.05 percent), Indonesia (2 percent), and Nigeria (1.91 percent). Together, these 10 countries occupy 50 percent of the global cropland, and they have 52 percent of the global population. Their combined cropland area increased by 2 percent between 2000 and 2015, compared to the substantial increase in population of 517 million (15.5 percent). Together, India, United States, China, and Russia encompass 36 percent of the total area. In the United States and Canada, from 2000 to 2015, cropland decreased by about 2 percent, whereas their populations increased by 14 and 13 percent, respectively. The additional food requirements in these 10 countries, which are caused by increased populations, as well as increasing nutritional demands, are met by production increases in existing cropland or through virtual food trade, or both. More than 18 countries, territories, departments, or regions had 60 percent or more of their geographic area as cropland: Republic of Moldova, San Marino, and Hungary had more than 80 percent of the country’s area as cropland; Denmark, Ukraine, Ireland, and Bangladesh, 70 to 80 percent; and Uruguay, Netherlands, United Kingdom, Spain, Lithuania, Poland, Gaza Strip, Czechia, Italy, India, and Azerbaijan, 60 to 70 percent. Europe and South Asia can be considered agricultural capitals of the world, on the basis of their percentages of geographic area as cropland. United States, China, and Russia, which all have high cropland areas, are ranked second, third, and fourth in the world; India is ranked first. However, the amount of cropland as a percentage of the country’s geographic area is relatively very low for United States (18.3 percent), China (17.7 percent), and Russia (9.5 percent), whereas it is 60.5 percent for India. Most African and South American countries, territories, departments, or regions have less than 15 percent of their geographic area as cropland. China and India together house 36 percent of the world’s population; however, between 2000 and 2015, the amount of China’s cropland area fell by 18.9 percent, owing to urban expansion and the abandonment of farmlands caused by demographic changes (that is, the movement of population from villages to cities). In contrast, China’s population grew by 10 percent. The amount of India’s cropland increased by 8.5 percent, whereas its population grew by 20 percent. This study showed that, out of the 10 leading cropland countries, Ukraine, Nigeria, Russia, and Indonesia showed an 18 to 31 percent increase in cropland areas, on the basis of GCEP30 by the year 2015, compared to 2000. Nigeria’s cropland area increased by 25 percent, and its population increased by 31 percent in the same period. In these countries, food security is maintained by cropland expansion, productivity increases, and virtual food trade. Nevertheless, this trend of increasing net-cropland area and productivity will likely become difficult to maintain, owing to diminishing arable lands and plateauing of 50 years of continual yield increases, requiring policymakers to explore novel and data-supported approaches to solving future food security issues. The GCEP30 product, which can be browsed at full resolution at www.croplands.org , has been released for public download and use through U.S. Geological Survey (USGS)–National Aeronautics and Space Administration (NASA) Land Processes Distributed Active Archive Center (see https://lpdaac.usgs.gov/news/release-of-gfsad-30-meter-cropland-extent-products/ ).

Professional Paper

A database of natural monthly streamflow estimates from 1950 to 2015 for the conterminous United States

Quantifying and understanding the natural streamflow regime, defined as expected streamflow that would occur in the absence of anthropogenic modification to the hydrologic system, is critically important for the development of management strategies aimed at protecting aquatic ecosystems. Water balance models have been applied frequently to estimate natural flows, but are limited in the number of predictor variables that can be included. Here, a statistical machine learning technique — random forest modeling — was applied to estimate natural flows at a monthly time‐step from 1950 to 2015 for >2.5 million stream reaches in the conterminous United States (U.S.) using 200 potential predictor variables. We describe the development and documentation of this dataset and assess model performance. Model fit statistics (mean Nash–Sutcliffe efficiency = 0.85; observed/expected ratio = 0.94) indicate good correspondence between predicted and observed flows at nearly 2,000 streamgages. As an example application of the dataset, the observed streamflow record at a site prior to and after the construction of an upstream reservoir was compared with estimated natural flows to demonstrate the magnitude of seasonal depletions in streamflow due to the reservoir. This dataset can be applied to quantify natural and anthropogenic processes contributing to streamflow depletion or augmentation, and assess associated ecological effects.

Journal of the American Water Resources Associatio

Modelling gully-erosion susceptibility in a semi-arid region, Iran: Investigation of applicability of certainty factor and maximum entropy models

Gully erosion susceptibility mapping is a fundamental tool for land-use planning aimed at mitigating land degradation. However, the capabilities of some state-of-the-art data-mining models for developing accurate maps of gully erosion susceptibility have not yet been fully investigated. This study assessed and compared the performance of two different types of data-mining models for accurately mapping gully erosion susceptibility at a regional scale in Chavar, Ilam, Iran. The two methods evaluated were: Certainty Factor (CF), a bivariate statistical model; and Maximum Entropy (ME), an advanced machine learning model. Several geographic and environmental factors that can contribute to gully erosion were considered as predictor variables of gully erosion susceptibility. Based on an existing differential GPS survey inventory of gully erosion, a total of 63 eroded gullies were spatially randomly split in a 70:30 ratio for use in model calibration and validation, respectively. Accuracy assessments completed with the receiver operating characteristic curve method showed that the ME-based regional gully susceptibility map has an area under the curve (AUC) value of 88.6% whereas the CF-based map has an AUC of 81.8%. According to jackknife tests that were used to investigate the relative importance of predictor variables, aspect, distance to river, lithology and land use are the most influential factors for the spatial distribution of gully erosion susceptibility in this region of Iran. The gully erosion susceptibility maps produced in this study could be useful tools for land managers and engineers tasked with road development, urbanization and other future development.

Science of the Total Environment

Analyzing multi-year nitrate concentration evolution in Alabama aquatic systems using a machine learning model

Rising nitrate contamination in water systems poses significant risks to public health and ecosystem stability, necessitating advanced modeling to understand nitrate dynamics more accurately. This study applies the long short-term memory (LSTM) modeling to investigate the hydrologic and environmental factors influencing nitrate concentration dynamics in rivers and aquifers across the state of Alabama in the southeast of the United States. By integrating dynamic data such as streamflow and groundwater levels with static catchment attributes, the machine learning model identifies primary drivers of nitrate fluctuations, offering detailed insights into the complex interactions affecting multi-year nitrate concentrations in natural aquatic systems. In addition, a novel LSTM-based approach utilizes synthetic surface water nitrate data to predict groundwater nitrate levels, helping to address monitoring gaps in aquifers connected to these rivers. This method reveals potential correlations between surface water and groundwater nitrate dynamics, which is particularly meaningful given the lack of water quality observations in many aquifers. Field applications further show that, while the LSTM model effectively captures seasonal trends, limitations in representing extreme nitrate events suggest areas for further refinement. These findings contribute to data-driven water quality management, enhancing understanding of nitrate behavior in interconnected water systems.

Alabama

Digital mapping of ecological land units using a nationally scalable modeling framework

Ecological site descriptions (ESDs) and associated state-and-transition models (STMs) provide a nationally consistent classification and information system for defining ecological land units for management applications in the United States. Current spatial representations of ESDs, however, occur via soil mapping and are therefore confined to the spatial resolution used to map soils within a survey area. Land management decisions occur across a range of spatial scales and therefore require ecological information that spans similar scales. Digital mapping provides an approach for optimizing the spatial scale of modeling products to best serve decision makers and have the greatest impact in addressing land management concerns. Here, we present a spatial modeling framework for mapping ecological sites using machine learning algorithms, soil survey field observations, soil survey geographic databases, ecological site data, and a suite of remote sensing-based spatial covariates (e.g., hyper-temporal remote sensing, terrain attributes, climate data, land-cover, lithology). Based on the theoretical association between ecological sites and landscape biophysical properties, we hypothesized that the spatial distribution of ecological sites could be predicted using readily available geospatial data. This modeling approach was tested at two study areas within the western United States, representing 6.1 million ha on the Colorado Plateau and 7.5 million ha within the Chihuahuan Desert. Results show our approach was effective in mapping grouped ecological site classes (ESGs), with 10-fold cross-validation accuracies of 70% in the Colorado Plateau based on 1405 point observations across eight expertly-defined ESG classes and 79% in the Chihuahuan Desert based on 2589 point observations across nine expertly-defined ESG classes. Model accuracies were also evaluated using external-validation datasets; resulting in 56 and 44% correct classification for the Colorado Plateau and Chihuahuan Desert, respectively. National coverage of the training and covariate data used in this study provides opportunities for a consistent national-scale mapping effort of ecological sites.

Chihuahuan Desert, Colorado Plateau

Total uncertainty quantification in inverse solutions with deep learning surrogate models

We propose an approximate Bayesian method for quantifying the total uncertainty in inverse partial differential equation (PDE) solutions obtained with machine learning surrogate models, including operator learning models. The proposed method accounts for uncertainty in the observations, PDE, and surrogate models. First, we use the surrogate model to formulate a minimization problem in the reduced space for the maximum a posteriori (MAP) inverse solution. Then, we randomize the MAP objective function and obtain samples of the posterior distribution by minimizing different realizations of the objective function. We test the proposed framework by comparing it with the iterative ensemble smoother and deep ensembling methods for a nonlinear diffusion equation with an unknown space-dependent diffusion coefficient. Among other applications, this equation describes the flow of groundwater in an unconfined aquifer. Depending on the training dataset and ensemble sizes, the proposed method provides similar or more descriptive posteriors of the parameters and states than the iterative ensemble smoother method. Deep ensembling underestimates uncertainty and provides less-informative posteriors than the other two methods. Our results show that, despite inherent uncertainty, surrogate models can be used for parameter and state estimation as an alternative to the inverse methods relying on (more accurate) numerical PDE solvers.

Journal of Computational Physics