Geology ReportsSearch

FIND YOUR NEXT DISCOVERY

Results for “JGR Machine Learning and Computation”

Original records, connected by a shared subject.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

813 recordsLinked to original sources

SlideDetect: Spatio-temporal landslide detection using a three-dimensional convolutional neural network

Landslides pose a serious and ongoing threat to both human lives and infrastructure worldwide; therefore, it is of interest to predict where and when landslides are likely to occur. Advances in machine learning techniques have spurred numerous studies aimed at estimating relative landslide propensity, but are limited to spatial (as opposed to temporal) prediction due to the sparsity of landslide timing data. We address this data gap by training SlideDetect, a 3-dimensional convolutional neural network (3D CNN), to identify landslides based on their spatial and temporal occurrence within multitemporal image stacks. We use an inventory of landsides triggered by the 2018 Hokkaido earthquake and two years of monthly composite optical imagery spanning this event. The model can identify not only landslide location but also landslide date with an area under the precision-recall curve (PR-AUC) of 0.84. We further present a new standard for presenting PR curve results that explicitly compares model performance at different confidence thresholds, allowing for clearer model evaluation and comparison. Our new approach to constraining landslide timing paired with this more consistent and objective method for evaluating model performance shows considerable promise, and with further application and testing, SlideDetect could enhance the data availability and tools needed to advance landslide hazard and risk assessments.

JGR Machine Learning and Computation

Localization of spatiotemporally heterogeneous subsurface flows using autoencoder-based deep learning framework for time-lapse self-potential tomography

Self-potential (SP) monitoring has emerged as a valuable method for characterizing subsurface hydrogeological features and processes due to its sensitivity to fluid-induced electrokinetic effects. Despite advancements in SP inversion, challenges remain in imaging groundwater dynamics from SP activities due to complex hydrological settings and transient noise. In this study, a deep learning autoencoder (AE)-based framework is proposed for the spatiotemporal localization of subsurface fluid movement from time-lapse SP tomography. Temporal segments of time-lapse numerical inversions were first derived from long-term SP monitoring conducted from a floodplain site in Oak Ridge, Tennessee, known for active hyporheic exchange. Subsequently, AE models based on vision transformer (ViT), convolutional long short-term memory (ConvLSTM), convolutional neural network, and temporal convolutional network were individually trained and compared on the SP tomography segments for reconstruction performance. Finally, the reconstruction error over time serves as an anomaly score to identify moments of active SP variation, whereas spatial distributions of errors within these moments are analyzed to image and localize regions associated with anomalous subsurface fluid movement. The results demonstrate that ConvLSTM- and ViT-AE are most capable for the localization task with contrasting error distributions and consistent delineation of anomalies. Applying the method to both SP arrays parallel and perpendicular to the stream produced consistent anomaly zones near a fault or karst feature, validating the robustness and generalization of the approach. These results demonstrate the potential of the proposed framework as a scalable and interpretable tool for spatiotemporal analysis of subsurface flow dynamics in complex hydrogeological systems.

Tennessee

Extracting data from maps: Lessons learned from the artificial intelligence for critical mineral assessment competition

The U.S. Geological Survey (USGS), Defense Advanced Projects Research Agency (DARPA), NASA Jet Propulsion Laboratory (JPL), and MITRE ran a 12-week machine learning competition aimed at accelerating development of AI tools for critical mineral assessments. The Artificial Intelligence for Critical Mineral Assessment Competition solicited innovative solutions for two challenges: 1) automated georeferencing of historical maps, and 2) automated feature extraction from historical maps. Competitors used a new dataset of historical map images to train, validate, and evaluate their models. Automated georeferencing pipelines attained a median root-mean square error of 1.1 km. Prompt-based extraction (i.e., with user input) of polygons, polylines, and points from geologic maps yielded median F1-scores of 0.77, 0.56, 0.35, respectively. Geologic maps pose numerous challenges for AI workflows because they vary significantly. However, despite its short duration, the competition yielded promising results that have since spurred further innovation in this area and led to the development of new AI tools to semi-automate key, time-consuming parts of the assessment workflow.

Applied Computing and Geosciences

Total uncertainty quantification in inverse solutions with deep learning surrogate models

We propose an approximate Bayesian method for quantifying the total uncertainty in inverse partial differential equation (PDE) solutions obtained with machine learning surrogate models, including operator learning models. The proposed method accounts for uncertainty in the observations, PDE, and surrogate models. First, we use the surrogate model to formulate a minimization problem in the reduced space for the maximum a posteriori (MAP) inverse solution. Then, we randomize the MAP objective function and obtain samples of the posterior distribution by minimizing different realizations of the objective function. We test the proposed framework by comparing it with the iterative ensemble smoother and deep ensembling methods for a nonlinear diffusion equation with an unknown space-dependent diffusion coefficient. Among other applications, this equation describes the flow of groundwater in an unconfined aquifer. Depending on the training dataset and ensemble sizes, the proposed method provides similar or more descriptive posteriors of the parameters and states than the iterative ensemble smoother method. Deep ensembling underestimates uncertainty and provides less-informative posteriors than the other two methods. Our results show that, despite inherent uncertainty, surrogate models can be used for parameter and state estimation as an alternative to the inverse methods relying on (more accurate) numerical PDE solvers.

Journal of Computational Physics

A glimpse into the future of tectonic tremor monitoring

Tectonic tremor is a weak, long-duration seismic signal often observed in subduction zones and on some other plate-bounding faults. Because of tremor's characteristically low amplitude (and low signal-to-noise) and lack of clear phase arrivals, detecting and locating tremor usually requires techniques distinct from those applied to typical earthquakes. Major advances in detection and understanding of tremor have derived in the past from a powerful combination of new data and new analysis techniques. In a recent study, Sagae et al. (2025, https://doi.org/10.1029/2025jb031348 ) exploit that combination again, developing a new machine-learning based workflow and applying it to the S-net cabled seismic network in the Japan trench offshore northern Honshu. Their approach, although complex, succeeds in detecting several times more tremor activity than earlier studies, resulting in new insights and providing a blueprint for similar approaches that could be applied elsewhere. As real-time earthquake monitoring adopts similar tools, it may present an opportunity to bring tremor monitoring into operational workflows. In turn, this could solidify tremor monitoring as a component of future operational earthquake forecasting.

JGR Solid Earth

Evaluation of daily stream temperature predictions (1979-2021) across the contiguous United States using a spatiotemporal aware machine learning algorithm

Stream temperature controls a variety of physical and biological processes that affect ecosystems, human health, and economic activities. We used 42 years (1979–2021) of data to predict daily summary statistics of stream temperature across >50,000 stream reaches in the contiguous United States using a recurrent graph convolution network. We comprehensively documented the performance – both across all reaches and by stream type (e.g., reservoir or groundwater influence) – as a baseline for future improvement. The model showed reach-level RMSE of <2 °C with 90 % prediction intervals that contain 90.7 % of observations. We also assessed how the model captured variability in ecologically relevant metrics (e.g., R 2 for annual 7-day maximum = 0.76; R 2 for days exceeding 25 °C = 0.75). This model does not outperform state-of-the-art machine learning efforts (e.g., RMSE ≤1.5 °C) due to a limited input set but does provide the most spatially complete modeling to date to support water availability assessments.

contiguous United States

Multiple machine-learning estimation of groundwater levels and trends for the regional Mississippi River Valley alluvial aquifer

The Mississippi River Valley alluvial aquifer provides irrigation, public, and domestic water supplies across the south-central United States. Declining groundwater levels require improved characterization of changing conditions. Traditional potentiometric-surface mapping does not use all available water-level data or quantify uncertainty. To address these limitations, we developed a data-driven multiple machine-learning (MML) framework delivered through two open-source R packages. The covMRVAgen1 software assembles covariates to 155,960 monthly groundwater levels from 57,695 wells; the mmlMRVAgen1 software trains Cubist and Random Forest models, blends them, and makes 1-kilometer gridded predictions of monthly potentiometric surfaces for the period January 1980–December 2022. The MML approach provides a methodological foundation for region-scale spatiotemporal groundwater prediction and uncertainty quantification, generating 90-percent prediction limits with appropriate empirical coverage. Model performance is acceptable, with a root-mean-square error of about 4.2 feet, standard deviation of 24.82 feet, and a normalized Nash–Sutcliffe efficiency of 0.973.

Arkansas, Illinois, Louisiana, Mississippi, Missou

The use of fluorite geochemistry and machine learning to identify critical mineral systems

Fluorite (CaF 2 ) is a potential pathfinder to critical mineral and rare earth element (REE) deposits but its application has been limited to a narrow range of mineralization types. I show that fluorite is a robust recorder of mineralization fertility by applying statistical and machine-learning methods to a new global fluorite geochemical database. Distinct median rare earth and trace element patterns are observed among deposit types and genetic environments. Fluorite associated with carbonatites and REE deposits are relatively enriched in Sr and have minimal Eu anomalies. These characteristics define new bivariate discrimination diagrams that correctly identify 78% of carbonatite-related fluorite and 88% of fluorite from REE deposits. Random forest classifiers were developed for a wide range of mineralization types and genetic settings. Trained solely on rare earth element patterns, these models achieve accuracies of 77–79%. Higher classification accuracies (up to 88–96%) are obtained when including elements such as Sr, highlighting the significance of trace elements for optimal fluorite classification. The recognition of diagnostic fluorite compositional fingerprints, particularly in REE-fertile systems, underscores its potential as a pathfinder and indicator for critical mineral exploration in F-bearing environments.

Mineralium Deposita

Review and synthesis of the applications of machine learning to coalbed methane recovery

Over the last 30 years, a substantial literature has evolved on the use of machine learning (ML) to assess, predict, and improve the efficiency of coalbed methane (CBM) recovery. In the United States, the production of CBM declined as shale gas production matured, but CBM continues to be an important energy resource in other parts of the world. ML applications that have the potential to improve CBM reservoir management and production forecasts, and to increase exploration and operational efficiency, are still of significant interest. The integration of geostatistical techniques into the CBM ML applications has been largely absent but represents an opportunity for improvement. The literature demonstrates the widespread interest in, and applicability of, ML algorithms applied to CBM problems, and that they continue to result in improvements in predictive performance. However, (1) much of the research is more academic than operational, (2) many results are based on simulations, or small or proprietary datasets, (3) ML performance information can be inconsistent and sometimes entirely omitted, (4) most methodologies are unique to the specific CBM situation and likely not generalizable, (5) no standard data repositories are available to directly compare the performance of competing algorithms, and (6) the spatial component is often omitted. Finally, relatively new ML protocols involving causality analysis and reinforced learning, as well as hybrid workflows combining both supervised and unsupervised learning, are anticipated to dominate the future investigations. Integration of geostatistical and geospatial analysis with ML should enhance performance.

Book chapter

Evaluating machine learning approaches to identify and predict oil and gas produced water lithium concentrations

Recently, the demand for battery-grade lithium has substantially increased, largely due to electrification of the transportation sector. The search for new lithium sources has turned to produced waters (frequently brines), a large-volume wastewater by-product of oil and gas extraction. Geochemical analysis indicates the presence of varying concentrations of lithium from produced water samples collected across the United States and represented in the U.S. Geological Survey’s National Produced Water Geochemical Database, as well as mixtures of Marcellus Shale produced water included in the Pennsylvania Department of Environmental Protection’s Oil and Gas Well Waste Reports. We first examined whether the geochemical signature of the lithium-bearing produced waters is sufficiently distinct so that machine learning (ML) can be used to correctly classify samples to the formation of origin. The produced water sample data used to assess classification accuracy were from the Marcellus Shale, Utica Shale and Point Pleasant Formation (Utica), and Smackover Formation oil and gas wells. Further, we evaluated the potential for ML to accurately classify Marcellus Shale produced water spatially (i.e., northeast versus southwest Pennsylvania). We then investigated whether ML algorithms applied to a suite of geochemical concentration data (i.e. Ba, Br, Cl, K, Mg, Sr) may be used to predict the lithium concentration of an unknown sample. Finally, we applied an estimated economic lithium grade cutoff of 150 milligrams per liter (mg/l) and assessed the utility of ML to predict whether a produced water sample would fall above or below the grade cutoff based on the suite of geochemical parameters. Four machine learning algorithms—Random Forest (RF), Gradient Boosting Trees (GBT), Extreme Boosting (XGBoost), and Deep Neural Networks (DNN) were assessed. This study successfully demonstrates that all four machine learning methods can precisely and accurately estimate lithium concentrations and geologic formation classification. The products of this study contribute to the growing body of knowledge aimed at expanding the lithium resource base within the United States.

Alabama, Arkansas, Florida, Georgia, Louisiana, Mi

Computational electromagnetic geophysics for groundwater system studies: A review on established practices and recent advances

Identifying effective solutions for locating groundwater resources and ensuring the quality of drinking water is increasingly urgent, given the challenges posed by climate change and population growth. This review investigates electromagnetic geophysical imaging techniques, in both time- and frequency-domain, that can provide valuable insights for groundwater assessment. We explore computational electromagnetic methods used to evaluate electromagnetic data and several recent hydrogeophysical case studies. As open-source frameworks for modeling electromagnetic geophysical problems become available, a broader range of researchers can interpret their data with computationally advanced software. We provide an overview of documented open-source codes for evaluating electromagnetic data and analyze various hydrological targets in relation to their electromagnetic surveying technique and the computational method applied. Furthermore, we evaluate the potential of advanced computational techniques, including three-dimensional modeling, non-deterministic inversion and machine learning, to couple geophysical with numerical groundwater modeling and apply it in groundwater system studies. Despite obstacles such as complexity and resource demands, our findings indicate that the quantification and integration of predictive uncertainties from both electromagnetic and hydrological data and simulations would significantly improve the reliability of hydrogeophysical models. This can lead to a deeper understanding of groundwater systems and improved management practices.

Journal of Hydrology

Modeling seawater intrusion along the Alabama coastline using physical and machine learning models to evaluate the effects of multiscale natural and anthropogenic stresses

Seawater intrusion threatens groundwater resources in coastal regions, including southern Baldwin County, Alabama, where the freshwater-saltwater interface dynamics remain poorly understood. To address this gap, this study uses combined physics-based and machine-learning models to quantify seawater intrusion caused by natural (storm surges) and anthropogenic (human activities) perturbations. The long short-term memory network and wavelet analysis were used to assess vertical aquifer vulnerabilities, revealing that the shallow part of the Coastal lowlands aquifer system (CL1) in the southern Baldwin County region is more susceptible to sea level rise and groundwater extraction than deeper aquifers. Based on these findings, a cross-sectional numerical model (physics approach) for the CL1 aquifer was developed to evaluate tidal and storm surge effects, using Tropical Storm Claudette (June 2021) as a case study. Results showed that tidal fluctuations had a minimal impact on the saltwater-freshwater interface location, whereas storm surges caused substantial inland movement, with effects lasting for nine months. The steady-state version of the three-dimensional (3D) physical model predicted seawater intrusion across the entire area, and convolutional neural network-based modeling further validated the model results. The 3D physical model was also applied to a smaller area to assess human impact on the saltwater interface due to two groundwater pumping scenarios (± 50% of the baseline pumping rate). Results revealed that a 50% increase in groundwater withdrawals caused seawater to advance ~ 320 m inland, whereas a 50% reduction led to a ~ 270-meter retreat. This study highlights the vulnerability of Alabama’s shallow coastal aquifers to seawater intrusion due to storm surges and human activities, and demonstrates that combining physics-based models with machine learning approaches can improve groundwater predictions, though its accuracy depends on the availability of site-specific data.

Alabama

Quantifying groundwater response and uncertainty in beaver-influenced mountainous floodplains using machine learning-based model calibration

Beavers ( Castor canadensis ) alter river corridor hydrology by creating ponds and inundating floodplains, and thereby improving surface water storage. However, the impact of inundation on groundwater, particularly in mountainous alluvial floodplains with permeable gravel/cobble layers overlain by a soil layer, remains uncertain. Numerical modeling across various floodplain structures considers topographic and sediment complexity and multidirectional flow, linking inundation to groundwater response. This study develops a model-data integration workflow to address uncertainty in groundwater response to beaver-induced inundations in a mountainous alluvial floodplain in the Upper Colorado River Basin. Uncertain factors include seasonal hydrologic dynamics, hydraulic conductivities, floodplain structures, and meteorological forcings. We employed an ensemble of groundwater models, based on geophysical and hydrologic data, with machine learning-based calibration using a neural density estimator. This allowed us to quantify the vertical flux from the soil layer to the permeable gravel bed, the down-valley underflow within the gravel bed, and their ratios. Results show a significant increase in the vertical flux relative to down-valley underflow, from 2% during dry pond periods to 20% during wet periods, serving as an analogy for conditions without and with beaver ponds. The study highlights the influence of floodplain structure on groundwater storage, water balance, and water quality impacted by beaver ponds. A thick gravel bed layer, with a large down-valley underflow, minimizes the effect of beaver-induced inundation on water quality. We emphasize the need for field-scale measurements of floodplain structure and improved characterization of evapotranspiration changes to reduce uncertainty in groundwater response.

Colorado

Surface variable‐based machine learning for scalable arsenic prediction in undersampled areas

In the United States, private wells are not federally regulated, and many households do not test for Arsenic (As). Chronic exposure is linked with multiple health outcomes, and risk can change sharply over short distances and with well depth. Coarse maps or sparse sampling often miss exceedances. Most existing models operate at ∼1 km resolution and use groundwater chemistry or detailed geologic logs, which limits their use in undersampled areas where improved guidance is most needed. We overcome these limitations by developing a machine learning model for Minnesota, USA, that predicts As exposure risk using only surficial variables from remote sensing and global data sets. Variables related to surface water hydrology and geomorphology are selected based on mechanistic links that control redox conditions and As mobilization. Local training was essential, and surficial geology variables that are more sensitive to local conditions were needed to maximize model accuracy. The resulting complete model was sufficiently sensitive to generate accurate and detailed risk maps and depth profiles of As concentrations above the 10 μg/L maximum contaminant level. Accuracy depended on local training data density. We identified a training data density of 0.07 wells/km 2 as a practical target for stable county-level performance. Maps of exceedance probabilities highlight priority areas for testing that are particularly important in rural communities that have received less sampling. These results support public health action by guiding where to install wells and where to test them, how much new sampling is needed, and where treatment outreach is most urgent.

Minnesota

HyFlood: A surrogate-model-based framework for compound coastal flooding

Compound coastal flooding is a major threat to low-lying coastal regions and is expected to intensify under future climate change projections. However, modeling the joint interaction of waves, storm surge, tides, and rainfall remains computationally demanding, limiting the development of fast and reliable forecast tools. Here we present HyFlood, a hybrid statistical-numerical downscaling framework capable of computing and mapping high-resolution compound flood hazards while substantially reducing the computational cost compared with fully process-based hydrodynamic modeling. HyFlood combines statistical sampling and selection algorithms with a cascade of reduced-complexity surrogate models that emulate nearshore wave transformation, surf-zone hydrodynamics, and coastal, fluvial, and pluvial flooding. The surrogate models employ machine-learning and regression algorithms applied to a low-dimensional representation of the flooding outputs, obtained through statistical dimensionality reduction. The framework is demonstrated in southern O'ahu, Hawai'i, a region exposed to elevated sea levels driven by tides, waves, and storm surge along with frequent precipitation-driven flash flooding. Validation of the surrogates against the physics-based model outputs demonstrates that HyFlood accurately reproduces daily maxima of spatially distributed flooding depths. This hybrid approach offers a scalable and efficient tool to better quantify how changes in flooding drivers translate into hazard and impact assessments, and to support compound-flood risk assessments and climate-change adaptation planning.

Hawaii

Efficient physics‐informed ground‐motion simulations with reduced‐order models: CyberShake implications and high‐resolution site terms for southern San Andreas fault earthquakes

Recent advances in Probabilistic Seismic Hazard Analysis (PSHA) leverage physics‐based ground‐motion simulations to estimate seismic hazard, such as the CyberShake project. However, computational costs quickly escalate when performing PSHA for numerous faults or sites and can become prohibitively expensive. To reduce computational demands, CyberShake uses reciprocity and interpolates physics‐informed corrections from simulations conducted at fewer locations, but the accuracy of these interpolations remains poorly quantified. To quantify the interpolation accuracy, we derive high‐resolution, frequency‐dependent site terms for southern California and compare them with interpolated site terms using the CyberShake approach. We accomplish this by performing a set of earthquake point‐source simulations distributed along the nonplanar fault geometry for the southern San Andreas fault (SSAF) extending from Bombay Beach to Lake Hughes. Using SeisSol, we simulate three minutes of viscoelastic seismic wave propagation for these sources and store the horizontal‐component Green’s functions for 480,000 sites. We then use a scientific machine learning approach based on interpolated proper orthogonal decomposition to construct an accurate reduced‐order model of the Green’s functions to efficiently predict effective amplitude spectra (EAS) for finite‐source rupture models of SSAF earthquakes. Using minimum curvature interpolation with tension, as used in CyberShake, we compare the interpolated site terms against our high‐resolution site terms. We identify local discrepancies with EAS differing by up to a factor of approximately three. Furthermore, we identify locations where unexpectedly high or low ground motions are missed when using the interpolated dataset for these earthquakes. We estimate that our approach may be used within CyberShake to reduce the time‐to‐solution by a factor of 336 for the entire earthquake rupture forecast. Our analysis of physics‐based site terms provides more insight into the seismic hazard due to SSAF ruptures and guides future developments by combining high‐performance computing and reduced‐order modeling techniques for PSHA.

California

False positives in the identification of dynamic earthquake triggering

Dynamic earthquake triggering is commonly identified through the temporal correlation between increased seismicity rates and global earthquakes that are possible triggering events. However, correlation does not imply causation. False positives may occur when unrelated seismicity rate changes coincidently occur at around the time of candidate triggers. We investigate the expected false positive rate in Southern California with global M ≥ 6 earthquakes as candidate triggers. We compute the false positive rate by applying the statistical tests used by DeSalvio and Fan (2023), https://doi.org/10.1029/2023jb026487 to synthetic earthquake catalogs with no real dynamic triggering. We find a false positive rate of ∼3.5%–8.5% when realistic earthquake clustering is present, consistent with the 95% confidence typically used in seismology. However, when this false positive rate is applied to the tens of thousands of spatial-temporal windows in Southern California tested in DeSalvio and Fan (2023), https://doi.org/10.1029/2023jb026487 , thousands of false positives are expected. The expected false positive occurrence is large enough to explain the observed apparent triggering following 70% of large global earthquakes (DeSalvio & Fan, 2023, https://doi.org/10.1029/2023jb026487 ), without requiring any true dynamic triggering. Aside from the known triggering from the nearby El Mayor-Cucapah, Mexico, earthquake, the spatial and temporal characteristics of the reported triggering are indistinguishable from random false positives. This implies that best practice for dynamic triggering studies that depend on temporal correlation is to estimate the false positive rate and investigate whether the observed apparent triggering is distinguishable from the correlations that may occur by chance.

JGR Solid Earth