Geology ReportsSearch

SEARCH · Geology Reports

Results for “Machine Learning with Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A review of machine learning applications to coastal sediment transport and morphodynamics

A range of computer science methods under the heading of machine learning (ML) enables the extraction of insight and quantitative relationships from multidimensional datasets. Here, we review some common ML methods and their application to studies of coastal morphodynamics and sediment transport. We examine aspects of ‘what’ and ‘why’ ML methods contribute, such as ‘what’ science problems ML tools have been used to address, ‘what’ was learned when using ML, and ‘why’ authors used ML methods. We find a variety of research questions have been addressed, ranging from small-scale predictions of sediment transport to larger-scale sand bar morphodynamics and coastal overwash on a developed island. We find various reasons justify the use of ML, including maximize predictability, emulation of model components, smooth and continuous nonlinear regression through data, and explicit inclusion of uncertainty. Overall the expanding use of ML has allowed for an expanding set of questions to be addressed. After reviewing the studies we outline a set of ‘best practices’ for coastal researchers using machine learning methods. Finally we suggest possible areas for future research, including the use of novel machine learning techniques and exploring ‘open data’ that is becoming increasingly available.

Earth-Science Reviews

Machine learning application to assess occurrence and saturations of methane hydrate in marine deposits offshore India

Artificial Neural Networks (ANN) were used to assess methane hydrate occurrence and saturation in marine sediments offshore India. The ANN analysis classifies the gas hydrate occurrence into three types: methane hydrate in pore space, methane hydrate in fractures, or no methane hydrate. Further, predicted saturation characterizes the volume of gas hydrate with respect to the available void volume. Log data collected at six wells, which were drilled during the India National Gas Hydrate Program Expedition 02 (NGHP-02), provided a combination of well log measurements that were used as input for machine learning (ML) models. Well log measurements included density, porosity, electrical resistivity, natural gamma radiation, and acoustic wave velocity. Combinations of well logs used in the ML models provide good overall balanced accuracy (0.79 to 0.86) for the prediction of the gas hydrate occurrence and good accuracy (0.68 to 0.92) for methane hydrate saturation prediction in the marine accumulations against reference data. The accuracy scores indicate that the ML models can successfully predict reservoir characteristics for marine methane hydrate deposits. The results indicate that the ML models can either augment physics-driven methods for assessing the occurrence and saturation of methane hydrate deposits or serve as an independent predictive tool for those characteristics.

Journal Interpretation

Review and synthesis of the applications of machine learning to coalbed methane recovery

Over the last 30 years, a substantial literature has evolved on the use of machine learning (ML) to assess, predict, and improve the efficiency of coalbed methane (CBM) recovery. In the United States, the production of CBM declined as shale gas production matured, but CBM continues to be an important energy resource in other parts of the world. ML applications that have the potential to improve CBM reservoir management and production forecasts, and to increase exploration and operational efficiency, are still of significant interest. The integration of geostatistical techniques into the CBM ML applications has been largely absent but represents an opportunity for improvement. The literature demonstrates the widespread interest in, and applicability of, ML algorithms applied to CBM problems, and that they continue to result in improvements in predictive performance. However, (1) much of the research is more academic than operational, (2) many results are based on simulations, or small or proprietary datasets, (3) ML performance information can be inconsistent and sometimes entirely omitted, (4) most methodologies are unique to the specific CBM situation and likely not generalizable, (5) no standard data repositories are available to directly compare the performance of competing algorithms, and (6) the spatial component is often omitted. Finally, relatively new ML protocols involving causality analysis and reinforced learning, as well as hybrid workflows combining both supervised and unsupervised learning, are anticipated to dominate the future investigations. Integration of geostatistical and geospatial analysis with ML should enhance performance.

Book chapter

Preliminary report on applications of machine learning techniques to the Nevada geothermal play fairway analysis

We are applying machine learning (ML) techniques, including training set augmentation and artificial neural networks, to mitigate key challenges in the Nevada play fairway project. The study area includes ~85 active geothermal systems as potential training sites and >12 geologic, geophysical, and geochemical features. The main goal is to develop an algorithmic approach to identify new geothermal systems in the Great Basin region. Major objectives include: 1) integrate ML techniques into the geothermal community; 2) develop open community datasets, whereby all play fairway and ML datasets and algorithms are publicly released and available for modification by various user groups; 3) identify data acquisition targets with high value for future work; 4) identify new signatures to detect blind geothermal systems; and 5) foster new capabilities for characterizing subsurface temperature and permeability. Initially, ML techniques are being applied to the same play fairway datasets and workflow. ML will then be applied to both enhanced and additional datasets, with modification of the PFA workflow to incorporate the new datasets. Finally, ML will be applied to define new workflows using the enhanced and additional datasets. An algorithmic approach that empirically learns to estimate weights of influence for diverse parameters can potentially scale and perform better than the play fairway analysis. Initial work on this project has involved 1) evaluating potential positive and negative training sites, 2) transformation of datasets into formats suitable for ML, and 3) initial development and testing of ML techniques.

Nevada

Chapter 12 - Explainable AI for understanding ML-derived vegetation products

Current machine learning applications and algorithms have developed promise to produce autonomous systems that automatically perceive, learn, predict, and act on their own. However, the effectiveness of these systems is limited by the machine's current inability to explain their decisions, algorithmic paths, and actions to human users. The purpose of this chapter is to apply explainable artificial intelligence (XAI) to black-box models using an example of the U.S. Geological Survey's LANDFIRE Existing Vegetation Type (EVT). This chapter also demonstrates the tools developed to assist scientists/analysts in understanding and trusting prediction outcomes of vegetation type that streamline development of the LANDFIRE EVT product.

Book chapter

What did they just say? Building a Rosetta stone for geoscience and machine learning

Modern advancements in science and engineering are built upon multidisciplinary projects that bring experts together from different fields. Within their respective disciplines, researchers rely on precise terminology for specific ideas, principles, methods, and theories. Hence, the potential for miscommunication is substantial, especially when common words have been adopted by one (or both) group(s) to represent very specific, precise, but, perhaps, different concepts. Under the best circumstances, misunderstanding key terms will lead toward a breakdown of efficiency. Under less optimal conditions, miscommunication will sow frustration, lead to errors, and inhibit scientific breakthroughs. Here, our research group of geoscientists and machine learning experts presents a process to help geoscientists understand the fundamentals of supervised learning by describing the general workflow (i.e., a conceptual pipeline) for supervised learning that must be understood by all the parties involved in a geoscience-machine learning endeavor. Terms critical for machine learning are introduced, defined, and used within the context of an overly simplified mock hydrological study to illustrate their appropriate usage, and then used again in the context of a published geothermal-machine learning study. These key terms are divided into two groups, which are 1) essential to the field of machine learning but are predominantly absent in geoscience or 2) homonyms (i.e., words with the same spelling or pronunciation but with different meanings) between the fields. Lastly, we discuss a few other important homonyms that were not introduced in the general workflow but arise regularly in machine learning applications

Conference Paper

Physics-guided machine learning from simulation data: An application in modeling lake and river systems

This paper proposes a new physics-guided machine learning approach that incorporates the scientific knowledge in physics-based models into machine learning models. Physics-based models are widely used to study dynamical systems in a variety of scientific and engineering problems. Although they are built based on general physical laws that govern the relations from input to output variables, these models often produce biased simulations due to inaccurate parameterizations or approximations used to represent the true physics. In this paper, we aim to build a new data-driven framework to monitor dynamical systems by extracting general scientific knowledge embodied in simulation data generated by the physics-based model. To handle the bias in simulation data caused by imperfect parameterization, we propose to extract general physical relations jointly from multiple sets of simulations generated by a physics-based model under different physical parameters. In particular, we develop a spatio-temporal network architecture that uses its gating variables to capture the variation of physical parameters. We initialize this model using a pre-training strategy that helps discover common physical patterns shared by different sets of simulation data. Then we fine-tune it using limited observation data via a contrastive learning process. By leveraging the complementary strength of machine learning and domain knowledge, our method has been shown to produce accurate predictions, use less training samples and generalize to out-of-sample scenarios. We further show that the method can provide insights about the variation of physical parameters over space and time in two domain applications: predicting temperature in streams and predicting temperature in lakes.

Conference Paper

Machine learning and data augmentation approach for identification of rare earth element potential in Indiana Coals, USA

Rare earth elements and yttrium (REYs) are critical elements and valuable commodities due to their limited availability and high demand in a wide range of applications and especially in high-technology products. The increased demand and geopolitical pressures motivate the search for alternative sources of REYs, and coal, coal waste, and coal ash are considered as new sources for these critical elements. This research evaluates the REY potential of coals from Indiana (USA). However, although coal data revealed REY potential, it suffered from sparse samples with complete REY measurements. Therefore, we explore the applicability of machine learning (ML) models and data augmentation techniques to demonstrate their applicability to evaluate REY potential in Indiana, and other areas in coal basins, using selected coal parameters (Al2O3, Fe2O3, C, Ash, S, P, Mo, Zn, and As contents) as covariates (indicators). Due to the relatively small sample size with complete REY data in the Indiana Coal Database, two data augmentation techniques (Random Over-Sampling Examples and Synthetic Minority Over-Sampling Technique) were used. Four machine learning algorithms (linear discriminate analysis, support vector machine, random forest, and artificial neural networks) were applied for modeling REY potential as a classification problem. The results show that application of Synthetic Minority Over-Sampling Technique prior to development of the support vector machine (SVM) models generated the best REY classification with an accuracy of 95%. The encouraging results based on Indiana coal data may suggest that a similar approach can be used for other coal basins for screening the locations with REY potential. Those locations then can be targeted for more detailed geochemical surveys to identify most promising areas and evaluate overall REY resources.

Indiana

Machine learning for natural resource assessment: An application to the blind geothermal systems of Nevada

A study is underway to apply machine learning methods to evaluate natural resource potential. In particular, we are considering the search for blind geothermal systems in Nevada. Beginning with the data and experience from the previous Nevada play fairway analysis project, we are building models in TensorFlow/Keras and gaining experience toward predicting the geothermal resource potential as a probability map. During the first year of this project we have encountered several issues particular to using geological and geophysical data sets with these tools. Through an illustrative example we develop a promising workflow for future use as more data become available and are analyzed.

Nevada

Exploring the uncertainty of machine learning models and geostatistical mapping of rare earth element potential in Indiana coals, USA

Rare earth elements and yttrium (REEs) have a wide range of applications in high- and low-carbon technologies. The strategic significance of REEs has grown due to their expanding applications in manufacturing industries and the constrained availability of these essential resources. This research explores the applicability of machine learning models and their uncertainty for assessing the REE potential in coal beds using various coal parameters as inputs. The work focuses on developing a predictive model based on geological variables, excluding considerations related to potential shifts in the commodities market. The Indiana Coal Quality Database was used as the data source. The promising and unpromising indicators derived from the outlook coefficient of samples from the database were used as the REE potential indicator for machine learning classification models. The filter-based approach with bootstrap was used to evaluate the importance of the coal parameters and their prediction uncertainties. Four machine learning methods (linear discriminant analysis (LDA), random forest (RF), support vector machine (SVM), and artificial neural networks (ANN), a data balancing and augmentation approach (Synthetic Minority Over-sampling Technique), and bootstrap resampling techniques were used for building the models and evaluating their prediction capabilities under uncertainty. It was determined that the SVM bootstrap model with ten-times balanced and augmented data provided superior results compared with other models. Finally, stochastic spatial maps of the REE potential within the coal basin were generated using sequential indicator simulation. The spatial maps of the REE potential showed that a 29% area of the Indiana section of the Illinois coal basin has economic potential of REEs, with 90% confidence.

Indiana

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. Finally, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

Nevada

Machine learning provides reconnaissance-type estimates of carbon dioxide storage resources in oil and gas reservoirs

Oil and gas reservoirs represent suitable containers to sequester carbon dioxide (CO 2 ) in a supercritical state because they are accessible, reservoir properties are known, and they previously contained stored buoyant fluids. However, planners must quantify the relative magnitude of the CO 2 storage resource in these reservoirs to formulate a comprehensive strategy for CO 2 mitigation. Even reconnaissance-type estimates of CO 2 storage resources of known oil and gas reservoirs may require complicated calculations involving 1) estimates of recoverable oil and gas, 2) reservoir properties (depth, temperature, pressure, etc.), and 3) the physical qualities of the retained fluids. We demonstrate the application of machine learning (ML) algorithms to bypass these computations to yield more rapid estimates of CO 2 storage resources in reservoirs capable of hosting CO 2 in a supercritical state. ML algorithms are computationally efficient because they do not impose the strong assumptions on the data-generating process that standard statistical or engineering procedures require. Further, ML algorithms can capture highly complex, particularly nonlinear, relationships among predictor variables. We demonstrate the application of four different ML algorithms using data from onshore and offshore oil and gas reservoirs in Europe, and show they perform well when predictions are compared to engineering estimates. The proposed methods and models provide an effective and novel way to more rapidly and directly determine the subsurface CO 2 storage capacity of oil and gas reservoirs around the world, information that operators, researchers, and policymakers alike require to meet energy transition and decarbonization goals.

Frontiers in Enviornmental Science

Physics-guided machine learning for scientific discovery: An application in simulating lake temperature profiles

Physics-based models are often used to study engineering and environmental systems. The ability to model these systems is the key to achieving our future environmental sustainability and improving the quality of human life. This article focuses on simulating lake water temperature, which is critical for understanding the impact of changing climate on aquatic ecosystems and assisting in aquatic resource management decisions. General Lake Model (GLM) is a state-of-the-art physics-based model used for addressing such problems. However, like other physics-based models used for studying scientific and engineering systems, it has several well-known limitations due to simplified representations of the physical processes being modeled or challenges in selecting appropriate parameters. While state-of-the-art machine learning models can sometimes outperform physics-based models given ample amount of training data, they can produce results that are physically inconsistent. This article proposes a physics-guided recurrent neural network model (PGRNN) that combines RNNs and physics-based models to leverage their complementary strengths and improves the modeling of physical processes. Specifically, we show that a PGRNN can improve prediction accuracy over that of physics-based models (by over 20% even with very little training data), while generating outputs consistent with physical laws. An important aspect of our PGRNN approach lies in its ability to incorporate the knowledge encoded in physics-based models. This allows training the PGRNN model using very few true observed data while also ensuring high prediction accuracy. Although we present and evaluate this methodology in the context of modeling the dynamics of temperature in lakes, it is applicable more widely to a range of scientific and engineering disciplines where physics-based (also known as mechanistic) models are used.

ACM/IMS Transactions on Data Science

How machine learning can improve predictions and provide insight into fluvial sediment transport in Minnesota

Understanding fluvial sediment transport is critical to addressing many environmental concerns such as exacerbated flooding, degradation of aquatic habitat, excess nutrients, and the economic challenges of restoring aquatic systems. However, fluvial sediment transport is difficult to understand because of the multitude of factors controlling the potential sources, delivery, mechanics, and storage of sediment in aquatic systems. While physical fluvial sediment samples are an integral part of developing solutions for these environmental concerns, samples cannot be collected at every river and time of interest. Therefore, accurate and cost-effective estimates of sediment loading are needed to manage riverine sediment transport at a multitude of scales (Ellison et al. 2016); also needed are methods to estimate sediment transport at sites where little or no physical samples have been collected (Gray & Simes 2008). The application of machine learning (ML) approaches to estimate sediment transport has grown over the past two decades (Afan et al. 2016). ML used in sediment transport research has shown multiple benefits over traditional approaches, such as increased prediction accuracy, the ability to learn complex linear and non-linear relations amongst the dataset and providing the ability to interpret these complex relations with important features used in the model (Cisty et al. 2021; Francke et al. 2008; Khan et al. 2021; Zounemat-Kermani et al. 2020; Cutler et al. 2007).

Minnesota

Predicting redox-sensitive contaminant concentrations in groundwater using random forest classification

Machine learning techniques were applied to a large (n > 10,000) compliance monitoring database to predict the occurrence of several redox-active constituents in groundwater across a large watershed. Specifically, random forest classification was used to determine the probabilities of detecting elevated concentrations of nitrate, iron, and arsenic in the Fox, Wolf, Peshtigo, and surrounding watersheds in northeastern Wisconsin. Random forest classification is well suited to describe the nonlinear relationships observed among several explanatory variables and the predicted probabilities of elevated concentrations of nitrate, iron, and arsenic. Maps of the probability of elevated nitrate, iron, and arsenic can be used to assess groundwater vulnerability and the vulnerability of streams to contaminants derived from groundwater. Processes responsible for elevated concentrations are elucidated using partial dependence plots. For example, an increase in the probability of elevated iron and arsenic occurred when well depths coincided with the glacial/bedrock interface, suggesting a bedrock source for these constituents. Furthermore, groundwater in contact with Ordovician bedrock has a higher likelihood of elevated iron concentrations, which supports the hypothesis that groundwater liberates iron from a sulfide-bearing secondary cement horizon of Ordovician age. Application of machine learning techniques to existing compliance monitoring data offers an opportunity to broadly assess aquifer and stream vulnerability at regional and national scales and to better understand geochemical processes responsible for observed conditions.

Water Resources Research

Visualization of petroleum exploration maturity for six petroleum provinces outside the United States and Canada

Outside the United States and Canada, most of the world’s supplies of oil and natural gas are recovered from conventional (or discrete) oil and gas accumulations. This type of hydrocarbon accumulation remains a target for exploration. In this report, exploration and discovery data are used to visually assist in describing the exploration maturity of selected petroleum provinces with respect to conventional oil and natural gas accumulations. The specific provinces are the Campos Basin (Brazil), the Santos Basin (Brazil), the North Sea Graben (northwestern Europe), the Middle Magdelena Basin (Colombia), the Sirte Basin (Libya), and the Kutei Basin (Indonesia). For each province, discovery data and well data through October 2019 are reported; from these data, depth distributions of the oil in oil fields and natural gas in gas fields were computed. The concepts of delineated prospective area and explored area include elements of geographic spatial information and statistical data analytics. Graphs showing dynamic growth of discoveries that are tied to the delineated prospective area provide a means of grading prospective area. Visualizations put the results of exploration in the context of geographic and geologic features of the play or basin and can be a tool to assist geologists with the appraisal of the number and sizes of undiscovered petroleum accumulations. Visualizations of exploration drilling and discoveries can (1) assist in conceptualizing a geologic model of the basin, (2) highlight relations among discovered accumulations in different plays or assessment units within the basin, and (3) allow the geologist to identify the missing information needed to complete the geologic model of a basin. Further, if visualization attributes can be quantified, they may be used for formulating quantitative models that predict numbers and sizes of undiscovered oil and gas accumulations. Such modeling approaches include discovery process models, Bayesian network models that characterize play or assessment unit dependencies, and innovative applications of machine learning to complement standard geologic assessments. The purpose of this report is to show how visualizations can further the understanding of exploration maturity for the six selected petroleum provinces. It also shows how the geologic framework, geologic data, and drilling and discovery trends can give context to the interpretation of the visualizations that lead to appraisal of exploration maturity.

Campos Basin, Kutei Basin, Middle Magdelena Basin,

Methods to evaluate and improve the modeling of rupture directivity in assessment of seismic hazard

In recent years, there have been several advancements related to the modelling of near-source effects of earthquake rupture on strong ground shaking, leading to an improved characterization of ground motions and resulting seismic hazard. Some of these modifications have stemmed from physics-based numerical modelling of the earthquake rupture process, using physics-based dynamic rupture simulations. These contributions have led to a better understanding of how fault rupture characteristics, geometry, and the style of faulting can interact with the hypocenter-dependence on the path from source to site that may ultimately guide the development of seismic directivity models. Moving forward, the application of modern techniques can be used to incorporate these source characteristics and near-fault ground motion behavior that contribute to the azimuthally varying effects that result in rupture directivity. One example is the application of machine learning methods to support more automated integration of new predictor variables in model development and open more evaluation opportunities to access residuals. Here, we utilize several techniques to take advantage of the plethora of synthetic data and its ability to supplement preexisting trends observed in data. We showcase two examples of how models can be either developed, expanded upon, or constrained using artificial neural network model (ANNs). We evaluate the performance of the ANN with existing methods, comparing misfit, potential limitations, and ability to continue to improve upon these methods in the future. One approach uses a set of simulations with corresponding synthetic ground motions from the Southern California Earthquake Center (SCEC) CyberShake study to develop a ground motion model adapted to incorporate seismic directivity information using an ANN. This large database (TBs) enables us to train the model to capture magnitude, period, and distance variations and how these parameters relate to amplification from hypocenters located along finite-faults. In some cases, there is reduced misfit from better representing source features that aren’t included in base ground motion models that neglect hypocenter location (e.g. azimuthal variation, source-to-site terms). Another ANN method uses a shallow-layered neural network model to better fit a hypocenter-independent model. This method adjusts the median and aleatory variability to account for the averaged impact of various hypocenter distributions to fit the underlying directivity adjustment model. This method serves as a template to apply to other directivity models, improving computational efficiency and more readily enabling integration in hazard codes.

California

Model-based assessment and mapping of total phosphorus enrichment in rivers with sparse reference data

Water nutrient management efforts are frequently coordinated across thousands of water bodies, leading to a need for spatially extensive information to facilitate decision making. Here we explore potential applications of a machine learning model of river low-flow total phosphorus (TP) concentrations to support landscape nutrient management. The model was trained, validated, and then applied for all rivers of Michigan, USA to identify potential drivers of nutrient variation, predict alteration in nutrient concentrations from minimally disturbed conditions, and explore reach specific sensitivity to riparian agricultural change. A boosted regression tree model of low-flow TP concentrations trained on natural and anthropogenic landscape predictors accounted for 53 % of variation in cross-validation data, had good accuracy, little bias, and plausible relationships between predictors and response. Percent riparian agricultural cover accounted for the greatest root mean square error reduction in the modeled response (33.2 %), followed by riparian soil permeability (12.9 %), watershed slope (9.6 %), and percent urban cover (9.6 %). An apparent non-linear relationship between TP concentrations and percent riparian agricultural cover suggested steep positive increases in stream TP concentrations between 10 and 30 % upstream riparian agricultural cover. Predicted minimally disturbed TP concentrations were spatially variable and ranged from 7.0 to 48.5 μg l −1 , with the highest concentrations in watersheds draining low-permeability lake plain soils. Comparison of minimally disturbed predictions to those from the early 2000s suggested that much of northern Michigan existed close to the reference condition, while lower Michigan streams were often substantially enriched. Our predicted values of minimally disturbed condition generally agree with previous studies but offer greater geographic specificity. Expanded application of machine learning modeling with landscape predictor data have great potential to inform large scale strategy development in landscapes with sparse reference data.

Michigan