Geology ReportsSearch

Geology topics

Cary Ruth Lindsey

Publications and source records attributed to Cary Ruth Lindsey.

6 recordsLinked to original sources

When less is more: How increasing the complexity of machine learning strategies for geothermal energy assessments may not lead toward better estimates

Previous moderate- and high-temperature geothermal resource assessments of the western United States utilized data-driven methods and expert decisions to estimate resource favorability. Although expert decisions can add confidence to the modeling process by ensuring reasonable models are employed, expert decisions also introduce human and, thereby, model bias. This bias can present a source of error that reduces the predictive performance of the models and confidence in the resulting resource estimates. Our study aims to develop robust data-driven methods with the goals of reducing bias and improving predictive ability. We present and compare nine favorability maps for geothermal resources in the western United States using data from the U.S. Geological Survey's 2008 geothermal resource assessment. Two favorability maps are created using the expert decision-dependent methods from the 2008 assessment ( i.e., weight-of-evidence and logistic regression). With the same data, we then create six different favorability maps using logistic regression (without underlying expert decisions), XGBoost, and support-vector machines paired with two training strategies. The training strategies are customized to address the inherent challenges of applying machine learning to the geothermal training data, which have no negative examples and severe class imbalance. We also create another favorability map using an artificial neural network. We demonstrate that modern machine learning approaches can improve upon systems built with expert decisions. We also find that XGBoost, a non-linear algorithm, produces greater agreement with the 2008 results than linear logistic regression without expert decisions, because the expert decisions in the 2008 assessment rendered the otherwise linear approaches non-linear despite the fact that the 2008 assessment used only linear methods. The F1 scores for all approaches appear low (F1 score < 0.10), do not improve with increasing model complexity, and, therefore, indicate the fundamental limitations of the input features ( i.e., training data). Until improved feature data are incorporated into the assessment process, simple non-linear algorithms ( e.g., XGBoost) perform equally well or better than more complex methods ( e.g., artificial neural networks) and remain easier to interpret.

Geothermics

New maps of conductive heat flow in the Great Basin, USA: Separating conductive and convective influences

Geothermal well data from Southern Methodist University and the U.S. Geological Survey (USGS) were used to create maps of estimated background conductive heat flow across the Great Basin region of the western United States. These heat flow maps were generated as part of the USGS hydrothermal and Enhanced Geothermal Systems resource assessment process, and the creation process seeks to remove the influence of hydrothermal convection from the predictions of the background conductive heat flow. The heat flow maps were constructed using a custom-developed iterative process using weighted regression, in which convectively influenced outliers were de-emphasized by assigning lower weights to measurements with heat flow values further from the estimated local trend (e.g., local convective influence). The local linear weighted regression algorithm is two-dimensional locally estimated scatterplot smoothing where smoothness was controlled by varying the number of nearby wells used for each local interpolation. Three maps resulting from conductive heat flow models are detailed in this paper, highlighting the influence of measurement confidence. The three maps use either: measurements from all wells with equal weight (no confidence weights), or one of two different published categorization methods to de-emphasize low-quality measurements; one categorization method graded thermal gradient quality, the other categorization method graded thermal conductivity quality. Each map is an estimate of background conductive heat flow as a function of reported data quality, and a point coverage is also provided for all wells in the compiled dataset. The point coverage includes an important new attribute for geothermal wells: the residual, which can be interpreted as the departure of a well from the estimated background heat flow conditions, and the value of the residual may be useful in identifying the influence of fluids (hydrothermal or groundwater) on conductive heat flow. Of the three maps presented, the map that de-emphasized the impact of wells with low-quality thermal gradient measurements appears to perform best because it did not incorporate many of the wells in the Snake River Plain that do not penetrate the aquifer and are therefore very unlikely to reflect true conductive conditions.

Great Basin

Detrending Great Basin elevation to identify structural patterns for identifying geothermal favorability

Topography provides information about the structural controls of the Great Basin and therefore information that may be used to identify favorable structural settings for geothermal systems. The Nevada Machine Learning Project (NVML) tested the use of a digital elevation map (DEM) of topography as an input feature to predict geothermal system favorability. A recent study re-examines the NVML data, identifying the DEM as the most important feature, showing a broad uniform pattern of high-favorability in the lower-elevation west and low-favorability in the higher elevation east of their study area in north-central Nevada. This regional elevation trend conflicts with the geologic notion that local relative topography should be used to identify geologic structures associated with favorable structural settings for hydrothermal upflow. Specifically, local relative topography gives information about position in the mountains, in the valleys, or at the transitions between, aiding in identification of faults and fault intersections. As part of U.S. Geological Survey efforts to engineer features that are useful for predicting geothermal resources, we construct a detrended elevation map that emphasizes local relative topography and highlights features that geologists use for identifying geothermal systems (i.e., providing machine learning algorithms with features that may improve predictive skill by emphasizing the information used by geologists). Herein, we describe the removal of the regional trend in elevation to emphasize the basin-and-range scale structural features, creating detrended elevation maps. Regional elevation trends were estimated using a local linear regression and subtracted from the actual elevation using a 30-m DEM. In an effort to optimize the detrended surface, alternate versions were produced with different rates of smoothness resulting in three detrended elevation maps. The resulting elevation trend surfaces (a proxy for crustal thickness) are compared with conductive heat flow maps, and a general pattern was observed of a negative correlation between heat flow and regional elevation in many areas, indicating that thinner crust may be causing elevated heat flow in some areas and thicker crust may cause the observed heat flow lows. Because these detrended elevation maps emphasize geologic structure and relative displacement, these products may also be useful for other geologic research including mineral exploration, hydrologic research, and defining geologic provinces.

Geothermal Resources Council Transactions

Predicting geothermal favorability in the western United States by using machine learning: Addressing challenges and developing solutions

Previous moderate- and high-temperature geothermal resource assessments of the western United States utilized weight-of-evidence and logistic regression methods to estimate resource favorability, but these analyses relied upon some expert decisions. While expert decisions can add confidence to aspects of the modeling process by ensuring only reasonable models are employed, expert decisions also introduce human bias into assessments. This bias presents a source of error that may affect the performance of the models and resulting resource estimates. Our study aims to reduce expert input through robust data-driven analyses and better-suited data science techniques, with the goals of saving time, reducing bias, and improving predictive ability. We present six favorability maps for geothermal resources in the western United States created using two strategies applied to three modern machine learning algorithms (logistic regression, support-vector machines, and XGBoost). To provide a direct comparison to previous assessments, we use the same input data as the 2008 U.S. Geological Survey (USGS) conventional moderate- to high-temperature geothermal resource assessment. The six new favorability maps required far less expert decision-making, but broadly agree with the previous assessment. Despite the fact that the 2008 assessment results employed linear methods, the non-linear machine learning algorithms (i.e., support-vector machines and XGBoost) produced greater agreement with the previous assessment than the linear machine learning algorithm (i.e., logistic regression). It is not surprising that geothermal systems depend on non-linear combinations of features, and we postulate that the expert decisions during the 2008 assessment accounted for system non-linearities. Substantial challenges to applying machine learning algorithms to predict geothermal resource favorability include severe class imbalance (i.e., there are very few known geothermal systems compared to the large area considered), and while there are known geothermal systems (i.e., positive labels), all other sites have an unknown status (i.e., they are unlabeled), instead of receiving a negative label (i.e., the known/proven absence of a geothermal resource). We address both challenges through a custom undersampling strategy that can be used with any algorithm and then evaluated using F1 scores.

western United States

What did they just say? Building a Rosetta stone for geoscience and machine learning

Modern advancements in science and engineering are built upon multidisciplinary projects that bring experts together from different fields. Within their respective disciplines, researchers rely on precise terminology for specific ideas, principles, methods, and theories. Hence, the potential for miscommunication is substantial, especially when common words have been adopted by one (or both) group(s) to represent very specific, precise, but, perhaps, different concepts. Under the best circumstances, misunderstanding key terms will lead toward a breakdown of efficiency. Under less optimal conditions, miscommunication will sow frustration, lead to errors, and inhibit scientific breakthroughs. Here, our research group of geoscientists and machine learning experts presents a process to help geoscientists understand the fundamentals of supervised learning by describing the general workflow (i.e., a conceptual pipeline) for supervised learning that must be understood by all the parties involved in a geoscience-machine learning endeavor. Terms critical for machine learning are introduced, defined, and used within the context of an overly simplified mock hydrological study to illustrate their appropriate usage, and then used again in the context of a published geothermal-machine learning study. These key terms are divided into two groups, which are 1) essential to the field of machine learning but are predominantly absent in geoscience or 2) homonyms (i.e., words with the same spelling or pronunciation but with different meanings) between the fields. Lastly, we discuss a few other important homonyms that were not introduced in the general workflow but arise regularly in machine learning applications

Conference Paper

Exploring declustering methodology for addressing geothermal exploration bias

Geothermal resources assessments use data that are unevenly distributed in space, with more data collected in areas with known thermal features. To meet the assumptions for geostatistical modeling (e.g., variography and kriging) such as having a random sample representative of the population, declustering may be needed to correct for spatial sample bias. Several declustering methods exist and to understand how best to use these methods, we apply these to real data and samples of that data. The work described herein summarizes the application of cell-based declustering to shallow temperature data (~20 cm) collected in a survey across a thermal feature in the Lower Geyser Basin, Yellowstone National Park, Wyoming. The sample dataset is a regular grid (3-m spacing) of temperatures across a 72-m square area, providing a shallow, subsurface temperature dataset collected with minimal spatial bias (a few grid locations near a hot spring could not be sampled). To test the influence of sample clustering on geothermal estimates, this dense dataset is sub-sampled irregularly to evaluate bias on temperature estimation. Three sampling strategies were tested: a simple random sample, a stratified random sample, and a stratified biased random sample. The naive mean (before declustering) values for each dataset were compared to the post-declustering mean to evaluate the effectiveness of declustering on correcting the mean for spatial bias. For the limited number of sample datasets evaluated, we found that although cell-based declustering did partially correct the mean, some bias remained (i.e., the estimate was improved, but not fully corrected). It is possible that the procedure documented herein (applied here to only a few random samples) could be applied to many random samples, so that robust conclusions might be drawn (e.g., Is there always some remaining bias in declustered estimates? Does it depend on the number of sample points?). In particular, bias could be evaluated for persistency, and uncertainty could be evaluated.

Geothermal Resources Council Transactions