Geology Reports⌕ Search

SEARCH · Geology Reports

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Design and documentation of a Baltimore-Washington regional spatial database testbed for environmental model calibration and verification

Recent efforts by scientists and managers to inventory, map, and model impacts of human activities on the environment have focused on land transformation and urbanization processes. To test the efficacy of any single model, algorithm or procedure which defines land transformation processes a standard database calibration reference resource is required. Therefore, a set of georeferenced, spatially structured and well documented data sets has been designed for the Baltimore-Washington Region as a test and evaluation resource for the community of environmental modelers and global change scientists. Land transformation processes are being examined from a variety of perspectives and scales using a variety of indicator parameters and mensuration variables. Tools and techniques applied to land transformation assessments range from creation of simple population expansion maps to change detection calculations using remotely sensed satellite data. A variety of point and cell growth models have been applied to simulate the land transformation phenomenon. These activities have demonstrated the reality that urbanization and land transformation processes involve complex interacting variables. A team of scientists are expanding the efforts of the USGS Human Impacts on Land Transformation (HILT) project to build an Internet accessible "collaboratory" containing quality controlled spatially referenced calibration and validation databases. The Baltimore-Washington Regional Testbed provides for the calibration, verification, and validation for multiple scalar, temporal, thematic, and spectral assessments or models. This design and documentation procedures for creating the Baltimore-Washington Regional "Collaboratory" are presented in relation to its use for environmental modeling applications.

Maryland↗

Integrating magnetotellurics, soil gas geochemistry and structural analysis to identify hidden, high enthalpy, extensional geothermal systems

We applied magnetotellurics (MT), diagnostic structural affiliations, soil gas flux, and fluid geochemistry to assist in identifying hidden, high-enthalpy geothermal systems in extensional regimes of the U.S. Great Basin. We are specifically looking for high-angle, low-resistivity zones and dilatant geologic structures that can carry fluids from magmatic or high-grade metamorphic conditions in the deep crust upward to exploitable depths, and to verify the nature of the deep sources through soil gas and fluid compositions. The project was motivated by prior MT transect coverage of western and central Nevada centered upon the Dixie Valley producing geothermal system where such favorable indicators were first recognized. The high-angle MT structures are taken to be fluidized fault zones connecting deep magmatic/metamorphic activity with the geothermal system, but the concept required verification by testing at other systems. The project was set up with a two-phased organization. Phase I was carried out at the McGinness Hills system, central Nevada, where Ormat Inc flagship power facility is located and a considerable amount of pre-existing data were available. Resistivity models along MT transects also showed a strong low-resistivity upwelling originating from interpreted deep crustal magmatic underplating. Controlling structures on production as indicated by Ormat data and our new mapping were favorable to dilatancy, comprising an accommodation zone between major normal faults of opposing dip. A 3D MT survey and inversion confirmed the existence of the steep low-resistivity zone dipping ESE toward the deep crust and placed N-S bounds upon the feature. In cooperation with Ormat personnel, we sampled well fluids from production intervals for He isotope composition. Elevated 3He was verified through mass spectrometry analysis confirming a magmatic connection with the producing system. High CO2 soil gas flux including possibly metamorphic 13C and 14C component was measured over the area of dilatant structures. Hence, the triad of indicators posed above was confirmed in Phase I. Subsequently, Phase II of the project proceeded in the greenfield Kumiva-Blackrock Desert district of northwestern Nevada to see if a new system could be identified. Transect MT data also showed a low-resistivity upwelling originating from interpreted deep crustal magmatic underplating. An MT survey of 131 sites was imaged through 3D inversion using an in-house, DOE-supported finite element algorithm. Low resistivity upwellings that warranted follow up study occur under the flanks of the Seven Troughs Range, under Kumiva Playa immediately west of the Blue Wing Mountains, and under northern Granite Springs Valley. Structural assessment of the project area by Co-I J. Faulds at UNR provided numerous favorable Quaternary fault settings, which were correlated to the MT upwelling structures. Soil CO2 gas flux anomalies generally were not large but did show correlation with resistivity upwelling structure and favorable geological structures. Isotope analyses showed presence of possible inorganic/metamorphic 13C but 14C concentrations did not exceed background values. We view the initial concept of a confluence of low-resistivity upwelling, favorably dilatant 3D geological structure, and elevated soil gas flux including 13C component to be supported by the further evidence of this project although the indicators in the Phase II study were more diffuse. Mass balance calculations based upon 3He R/Ra values indicates that the proportion of magmatic fluids in a producing system is fairly low, 10-15% by volume. We suggest that the diagnostic MT geophysical structures denote zones of concentrated extensional deformation that increases permeability, potentially enabling a circulating upper crustal geothermal system, while at the same time connecting telltale deep component signatures to the upper crust. The northern Granite Springs Valley structure is receiving followup stu

Nevada↗

Object-based inversion of crosswell radar tomography data to monitor vegetable-oil injection experiment

Crosswell radar tomography methods can be used to dynamically image ground-water flow and mass transport associated with tracer tests, hydraulic tests, and natural physical processes. Dynamic imaging can be used to identify preferential flow paths and to help characterize complex aquifer heterogeneity. Unfortunately, because the raypath coverage of the interwell region is limited by the borehole geometry, the tomographic inverse problem is typically underdetermined, and tomograms may contain artifacts such as spurious blurring or streaking that confuse interpretation. We implement object-based inversion (using a constrained, non-linear, least-squares algorithm) as an alternative to pixel-based inversion approaches that utilize regularization (such as damping or smoothing criteria). Our approach requires pre- and post-injection travel-time data. Parameterization of the image plane comprises a small number of objects rather than a large number of pixels, resulting in an overdetermined problem that reduces the need for prior information. The nature and geometry of the objects are based on hydrologic insight into aquifer characteristics, the nature of the experiment, and the planned use of the geophysical results. The object-based inversion approach is demonstrated using synthetic and crosswell radar field data acquired during vegetable-oil injection experiments at a site in Fridley, Minnesota. The region where oil has displaced ground water is discretized as a stack of rectangles of variable horizontal extents. The inversion provides the geometry of the affected region and an estimate of the radar slowness change for each rectangle. Applying petrophysical models to these results and porosity from neutron logs, we estimate that the vegetable-oil emulsion saturation in various layers ranges from 60 to 90%. Further work is needed to assess the accuracy of the emulsion saturation estimates. Using synthetic- and field-data examples, the object-based inversion approach is shown to be an effective strategy for inverting crosswell radar tomography data acquired to monitor the emplacement of vegetable-oil emulsions. A principal advantage of object-based inversion is that it yields images that hydrologists and engineers can easily interpret and use for model calibration.

Conference Paper↗

Cross‐hole radar attenuation tomography using a frequency centroid down‐shift method: Consideration of non‐linear frequency dependence of EM wave attenuation

This paper presents a cross-hole radar attenuation tomography method based on analysis of the down-shift in the spectrum centroid frequency, and spectral broadening of the received radar signals. The method uses a parameter that combines centroid frequency down shift and variance increase for the projection function to construct the tomography algorithm. In comparison with other methods for estimating attenuation, the frequency down shift method is relatively insensitive to the effects of geometric spreading, antenna coupling, antenna radiation pattern,and instrument response, but the method requires the data to be broad-band so the frequency shift is easily measured. This method is well suited for difference tomography when electrically conductive tracers are used. The method was tested using cross-hole radar data acquired before and during a saline tracer injection experiment at the U.S.Geological Survey’s Fractured Rock Research Site at Mirror Lake, in Grafton County, New Hampshire. The attenuation-difference tomogram clearly outlines the location of the saline tracer within the tomography plane.

New Hampshire↗

Data collection network to support ecosystem forecasting for the Barataria Basin - Mississippi River domain

Ecosystem forecasting is limited by a number of uncertainties including inadequate initialization information, unknown boundary conditions, inaccurate model physics and atmospheric forcing functions, and inadequate algorithm development of geomorphic and ecological responses to hydrodynamic and geophysical processes. Monitoring can help reduce these uncertainties by providing numerical information on those variables that can improve simulation capabilities. A data collection network was designed for the Mississippi River, south of Tarbert Landing, Ms. to Venice, La., and for the inland and nearshore waters of Barataria Basin, La. The network considered existing monitoring efforts that effectively link oceanographic and water quality observations to system drivers to support ecosystem hindcasting, nowcasting, and forecasting capabilities. The design was driven by the diverse needs of the modeling community and utilized questionnaire surveys, workshops, and web- based applications to inventory past and present modeling applications, domains, and attributes. The use of existing monitoring stations, including variables measured and frequency of data collection, were also considered, as well as recommendations on where new monitoring stations should be located, what types of data should be collected, and at what temporal frequency. Compiled databases are presented as a series of GIS layers to represent overlap between existing and requested stations, and priorities are established for either maintaining or augmenting existing stations or deploying new stations. An optimization plan will be prepared from the results of this study that will provide a rationale for the ecosystem forecasting data collection network that includes a description of existing monitoring programs to support ecosystem forecasting and a set of justifiable locations for "megastations" that meet the needs of the modeling and assessment communities.

Louisiana↗

Habitat affinities and at-sea ranging behaviors among main Hawaiian Island seabirds: Breeding seabird telemetry, 2013–2016

Recent Hawaiʻi state clean energy policy mandates and federal interest in developing offshore renewable energy resources have prompted unsolicited lease requests for offshore wind energy infrastructure (OWEI) to be located in ocean waters off Hawaiʻi. This study describing at-sea ranging behaviors for five seabirds was intended to provide new information on Hawaiian breeding seabird distribution at sea, habitat utilization, and ranging behaviors within near-island waters and throughout outer continental shelf (OCS) waters surrounding the main Hawaiian Islands (MHI). We also estimate the percentage of time the five study species spent flying at altitudes equivalent to an expected rotor-swept-zone (RSZ; 30–194 m) for an offshore wind turbine and discuss potential collision risk from OWEI to the seabirds studied here. The MHI supports important seabird breeding populations and individual seabirds can now be equipped with a wide-variety of data loggers and location tracking devices that can provide complex, high-resolution information on movement patterns at sea. In this study, we used GPS loggers and temperature-depth-recorders (TDRs) to examine the at-sea distributions and ranging behaviors of five abundantly breeding species in the MHI: Red-tailed Tropicbird, Laysan Albatross, Wedge-tailed Shearwater, Brown Booby, and Red-footed Booby. We tracked these breeding seabirds from 14 different sites throughout the MHI; study colonies were located on the main islands of Maui, Oʻahu, Kauaʻi, and on associated islets. We used the Residence in Space and Time (RST) algorithm to classify behavior into resting, transiting, and searching/foraging (Torres et al. 2017). We used GPS altitude measurements to examine species-specific flight height and to estimate time spent flying in the RSZ. We mapped rediscretized tracking data for seabirds at each study colony according to behavioral class and trip type (when applicable) using kernel density estimates. During 2014–2016, we obtained GPS and TDR data from 59 and 34 Red-tailed Tropicbirds, respectively. Taken together, individuals revealed a bimodal (short- [~3 h, <100 km range] and long- [>3 d, >800 km range]) trip foraging strategy. While ranging at sea, we estimated that Red-tailed Tropicbirds spend 70.6% (95% confidence interval [CI] 70.1–71.0%) of flight time in the RSZ. TDR data for tropicbirds was noisy and we could not reliably identify dives. During 2014 and 2016, we obtained GPS data from 35 Laysan Albatrosses nesting on Kauaʻi and Oʻahu. Individuals during the mid- to late-chick rearing period engaged in a bimodal short- (<6 d, <400 km range) and long- (>6 d, >2,000 km range) trip foraging strategy. While ranging at sea, we estimated that Laysan Albatrosses spend 2.57% (95% CI 2.50–2.64%) of flight time in the RSZ. During 2013–2015, we obtained GPS and TDR data from 313 and 55 Wedge-tailed Shearwaters, respectively. Considering all the data together, individuals revealed a multi-modal trip duration foraging strategy consisting of intra-day (<24 h, <100 km range), short (<4 d, ~200 km range), and long (>4 d, ~100–400 km range) trips. While ranging at sea, we estimated that Wedge-tailed Shearwaters spend 5.20% (95% CI 5.13–5.27%) of flight time in the RSZ. Wedge-tailed Shearwaters dove to a mean (± SD) depth of 1.78 ± 1.35 m (median = 1.38 m); the deepest dive recorded was to 10.06 m. The mean dive duration for Wedge-tailed Shearwaters was 3.12 ± 3.44 s (median = 1.80 s). During 2014–2015, we obtained GPS and TDR data from 42 and 37 Brown Boobies, respectively. Almost all foraging trips (97%) were single-day trips and we did not detect any bimodality in the distribution of single-day trip durations. Brown Boobies foraged relatively close to their colony (<60 km range) and focused their at-sea use in nearshore, coastal waters off Kauaʻi and Niʻihau. While foraging at sea, we estimated that Brown Boobies spend 3.41% (95% CI 3.16–3.67%) of flight time in the RSZ

Hawaii↗

Development of a new open-source tool to map burned area and burn severity

Accurate and complete geospatial fire occurrence records are important in determining postfire effects, emissions, hazards, and fuel loading inventories. Currently, the Monitoring Trends in Burn Severity (MTBS) project maps the fire perimeter and burn severity of all large fires on public lands. Although the MTBS project maps a large proportion of the fire acreage, it maps a smaller proportion of the actual number of fires in the United States, thereby creating a data gap. To fill this data gap, fire scientists at the U.S. Geological Survey (USGS) Earth Resources Observation and Science Center (EROS; Sioux Falls, South Dakota) proposed creating an open-source Fire Mapping Tool (FMT; available at https://mtbs.gov/qgis-fire-mapping-tool) as part of a two-phase National Aeronautics and Space Administration (NASA) Applied Fire Science Program grant. Phase II developed the FMT to map burn perimeters and severity not included in the MTBS database. This paper will focus on Phase II and will explain the algorithms that enhance the FMT’s functionality, demonstrate fire mapping procedures, and provide an example comparison between MTBS analyst fire products and those mapped using the FMT. The overall goal in the production of the FMT was to provide a freely available tool that can be used to map fires anywhere in the world.

Washington↗

2023 Coastal master plan: Model improvement plan, ICM-wetlands, vegetation, and soil

As part of the model improvement effort for the 2023 Coastal Master Plan, the wetland processes captured by the morphology and vegetation models used during previous master plans were reevaluated to assess how Integrated Compartment Model (ICM) subroutines could be improved. This process considered technical reviews, comments, and suggested improvements provided by model developers, advisory groups, and other experts during previous master plan cycles. The availability of new data and information that could be used to make model improvements was also considered. In many cases, the team considered and tested multiple options or approaches. As a result of this effort, recommended improvements are provided here. The improvements recommended to be included in the 2023 Coastal Master Plan include: adjusting marsh collapse thresholds, refining organic matter accretion calculations, developing an unstructured grid for modeling vegetation, improving flotant marsh and forested wetlands algorithms, creating and applying an updated map of existing vegetation, adjusting model code, and updating the submerged aquatic vegetation (SAV) module. This report describes the team’s work through a series of 7 distinct activities to identify and test options for model improvements to ensure the updated ICM used for the 2023 Coastal Master Plan appropriately captures ecological and morphological processes observed in Coastal Louisiana. As appropriate, relevant literature and data are discussed. Test runs to evaluate how changes influence model outputs are also documented. A final list of recommended updates, taking into account consideration of all options and results from test runs, is summarized at the end of the report. A later report will describe the final ICM-LAVegMod and ICM-Morph subroutines for the 2023 Coastal Master Plan, detailing the updates that have been incorporated.

Louisiana↗

Exploring basin-scale relations and unsupervised classification to quantify and automate the definition of assessment units in USGS continuous oil and gas resource assessments

The U.S. Geological Survey (USGS) assesses potential for undiscovered, technically recoverable oil and gas resources in priority geologic provinces and quantifies resource volume estimates within subdivisions called assessment units (AUs). AU boundaries are defined by USGS geologists using quantitative and qualitative geologic information. Variables contained in IHS Markit’s well and production databases can quantify and/or function as proxies for many of the qualitative, boundary-defining variables. This research explores a new approach to determine AU boundaries and the potential to automate their definition, using data analytics and machine learning algorithms on key, qualitative variables within the IHS Markit databases. Well and production data from the U.S. onshore Gulf Coast region for the Upper Cretaceous Eagle Ford Group and Austin Chalk are used in this analysis because each is relatively geologically uniform in Texas and both have recently been assessed by the USGS. The Eagle Ford is an example of an in situ continuous oil and gas accumulation, and the overlying Austin Chalk is an example of a combined conventional and continuous resource, sourced from the underlying Eagle Ford. Wellspecific values were extracted or calculated from data in IHS Markit’s well and production databases for depth to top and base of the formations, formation thickness, bottom-hole temperature, temperature gradient, temperature at base of formation, cumulative oil and gas production values, barrels of oil equivalent, oil and gas gravities, mud weights from initial well test, depth pressure ratio, and excess pressure. A raster for each variable was interpolated using the natural neighbor technique from the spatial analyst toolbox in ArcGIS. Rasters were then transformed using minimum-maximum scaling, which rescales the distribution to the range of 0–1. Clustering was completed using the iso cluster unsupervised classification tool on the normalized rasters. Raster cell groupings from two to ten were explored, with initial results demonstrating that four to six classes return the most differentiable groups, with depth to formation, oil gravity, pressure, and temperature variables containing the greatest between-group differences. Modeled clusters have spatial similarities to the geologically defined AUs, with indication that temperature and pressure are the most fundamental to AU definition. Input from geologists will remain crucial for further dividing clusters and defining final AUs, since AUs are defined by both qualitative and quantitative information; however, this research documents promising cluster modeling results for the automation of initial AU definitions.

Louisiana, Mississippi, Texas↗

The formation, transport, and breakup of submerged oil-particle aggregates in Great Lakes riverine environments

The formation, transport, and resuspension of oil-particle aggregates (OPA) in freshwater environments are of much interest to oil spill responders and scientists, especially as transportation of light and heavy crude oils has substantially increased across river corridors and coasts in the Great Lakes Basin. The persistent sheening from accumulated OPA along 60 km of the Kalamazoo River in Michigan’s lower peninsula resulted in a lengthy and expensive cleanup for the 2010 Enbridge Line 6B pipeline rupture. The interaction of oil with river mineral sediment and organic matter and its long-term fate depend on the physical properties of the oil and particles as well as the environmental setting of river, its climate, morphology, currents and mixing opportunities. This research brief describes the expanded work conducted for the cleanup for the 2010 Enbridge Line 6B pipeline rupture and includes laboratory experiments of aggregate characteristics with Cold Lake Blend and a range of sediment particle sizes, addition of an OPA formation algorithm to an existing sediment contaminant transport model, and development of a simplified, particle-tracking based rapid response model of OPA formation, transport, and deposition. A description of formulas developed for mixing energy in rivers in terms of river properties is also included.

Michigan↗

Solution of water-table and anisotropic flow problems by using the strongly implicit procedure

The use of the strongly implicit procedure (SIP) with an additional iteration parameter, β , to scale the residual vector is advantageous to the solution of some ground-waterflow problems. For steady-state water-table problems plagued by excessive elimination of grid blocks during the iteration process, selection of β <1 can be effective in limiting the deletion of blocks to a reasonable number. Also, a linear problem characterized by large anisotropy and layers of contrasting hydraulic conductivity was solved more efficiently with β =1.5. Effective values of β are generally in the range 0< β <2 and are easily determined by trial. Use of a β parameter in the SIP algorithm provides an effective solution technique for a class of ground-water-flow problems that previously was burdened by significant computational difficulty.

Journal of Research of the U.S. Geological Survey↗

Automated derivation of hydrologic basin characteristics from digital elevation model data

Digital elevation model (DEM) data in a raster format can be used to automatically derive the drainage characteristics of an area. A procedure has been designed that is capable of operating on matrices of elevation data having no algorithmically imposed size limit, while performing within the resolution and accuracy tolerances of the DEM data. Each cell is processed as the center of a 3- by 3-cell spatial window in the raster elevation data. If a cell is a local minimum in comparison with two of its non-adjacent neighbors, it is labeled as a drainage cell. The linkages of the drainage cells within user-specified distance and elevation thresholds are established in a separate process. The products of these processing steps are digital masks of the drainage cells and the watershed basins, both in raster format. A drainage cell mask derived using this procedure is useful in computing slope values for a raster data base. Slope has traditionally been calculated for each cell by fitting a plane through the eight nearest cells. However, if the terrain represented by these cells is V-shaped, such as a gully, a plane does not fit well; in fact, the desired slope value is the slope along the bottom of the gully, regardless of the steepness of the gully sides. The automated drainage process will label such a cell as a drainage cell, and its slope can then be computed from the elevation values of neighboring drainage cells.

Conference Paper↗

Methods and applications in surface depression analysis

Gridded surface data sets are often incorporated into digital data bases, but extracting information from the data sets requires specialized raster processing techniques different from those historically used on remotely sensed and thematic data. Frequently, the information desired of a gridded surface is directly related to the topologic peaks and pits of the surface. A method for isolating these peaks and pits has been developed, and two examples of its application are presented. The perimeter of a pit feature is the highest-valued closed contour surrounding a minimum level. The method devised for finding all such contours is designed to operate on large raster surfaces. If the data are first inversely mapped, this algorithm will find surface peaks rather than pits. In one example the depressions, or pits, expressed in Digital Elevation Model data, are hydrologically significant potholes. Measurement of their storage capacity is the objective. The potholes are found and labelled as polygons; their watershed boundaries are found and attributes are computed. In the other example, geochemical surfaces, which were interpolated from chemical analyses of irregularly distributed stream sediment samples, were analyzed to determine the magnitude, morphology, and areal extent of peaks (geochemical anomalies).

Conference Paper↗

Exploring the potential for a fused Landsat-MODIS snow covered area product

Results from nine 3 x 3 km study areas in the Rocky Mountains of Colorado, USA demonstrate there is potential for using sporadically acquired Landsat images in combination with daily coarse resolution fractional snow covered area (SCA) images to produce daily high resolution binary SCA images. The results also highlight several challenges to implementing this type of approach. The approach described here consistently yields accurate results in locations with persistent winter and spring snow cover where ten or more partially snow covered images are available to populate the image database, but is less successful in areas with shallower or more ephemeral snow covers or when fewer images are available to populate the image database. This work represents a first step towards developing an algorithm to combine Landsat and MODIS data to produce daily 30 m resolution binary SCA images. Further research should focus on testing the accuracy of this approach across a range of landscape types and snow cover regimes, developing methods to improve prediction accuracy when snow cover is nearly complete or nearly absent, and developing methods to compensate for the effects of canopy cover on SCA retrievals.

Conference Paper↗

New maps of conductive heat flow in the Great Basin, USA: Separating conductive and convective influences

Geothermal well data from Southern Methodist University and the U.S. Geological Survey (USGS) were used to create maps of estimated background conductive heat flow across the Great Basin region of the western United States. These heat flow maps were generated as part of the USGS hydrothermal and Enhanced Geothermal Systems resource assessment process, and the creation process seeks to remove the influence of hydrothermal convection from the predictions of the background conductive heat flow. The heat flow maps were constructed using a custom-developed iterative process using weighted regression, in which convectively influenced outliers were de-emphasized by assigning lower weights to measurements with heat flow values further from the estimated local trend (e.g., local convective influence). The local linear weighted regression algorithm is two-dimensional locally estimated scatterplot smoothing where smoothness was controlled by varying the number of nearby wells used for each local interpolation. Three maps resulting from conductive heat flow models are detailed in this paper, highlighting the influence of measurement confidence. The three maps use either: measurements from all wells with equal weight (no confidence weights), or one of two different published categorization methods to de-emphasize low-quality measurements; one categorization method graded thermal gradient quality, the other categorization method graded thermal conductivity quality. Each map is an estimate of background conductive heat flow as a function of reported data quality, and a point coverage is also provided for all wells in the compiled dataset. The point coverage includes an important new attribute for geothermal wells: the residual, which can be interpreted as the departure of a well from the estimated background heat flow conditions, and the value of the residual may be useful in identifying the influence of fluids (hydrothermal or groundwater) on conductive heat flow. Of the three maps presented, the map that de-emphasized the impact of wells with low-quality thermal gradient measurements appears to perform best because it did not incorporate many of the wells in the Snake River Plain that do not penetrate the aquifer and are therefore very unlikely to reflect true conductive conditions.

Great Basin↗

Detrending Great Basin elevation to identify structural patterns for identifying geothermal favorability

Topography provides information about the structural controls of the Great Basin and therefore information that may be used to identify favorable structural settings for geothermal systems. The Nevada Machine Learning Project (NVML) tested the use of a digital elevation map (DEM) of topography as an input feature to predict geothermal system favorability. A recent study re-examines the NVML data, identifying the DEM as the most important feature, showing a broad uniform pattern of high-favorability in the lower-elevation west and low-favorability in the higher elevation east of their study area in north-central Nevada. This regional elevation trend conflicts with the geologic notion that local relative topography should be used to identify geologic structures associated with favorable structural settings for hydrothermal upflow. Specifically, local relative topography gives information about position in the mountains, in the valleys, or at the transitions between, aiding in identification of faults and fault intersections. As part of U.S. Geological Survey efforts to engineer features that are useful for predicting geothermal resources, we construct a detrended elevation map that emphasizes local relative topography and highlights features that geologists use for identifying geothermal systems (i.e., providing machine learning algorithms with features that may improve predictive skill by emphasizing the information used by geologists). Herein, we describe the removal of the regional trend in elevation to emphasize the basin-and-range scale structural features, creating detrended elevation maps. Regional elevation trends were estimated using a local linear regression and subtracted from the actual elevation using a 30-m DEM. In an effort to optimize the detrended surface, alternate versions were produced with different rates of smoothness resulting in three detrended elevation maps. The resulting elevation trend surfaces (a proxy for crustal thickness) are compared with conductive heat flow maps, and a general pattern was observed of a negative correlation between heat flow and regional elevation in many areas, indicating that thinner crust may be causing elevated heat flow in some areas and thicker crust may cause the observed heat flow lows. Because these detrended elevation maps emphasize geologic structure and relative displacement, these products may also be useful for other geologic research including mineral exploration, hydrologic research, and defining geologic provinces.

Geothermal Resources Council Transactions↗

Predicting large hydrothermal systems

We train five models using two machine learning (ML) regression algorithms (i.e., linear regression and XGBoost) to predict hydrothermal upflow in the Great Basin. Feature data are extracted from datasets supporting the INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems project (INGENIOUS). The label data (the reported convective signals) are extracted from measured thermal gradients in wells by comparing the total estimated heat flow at the wells to the modeled background conductive heat flow. That is, the reported convective signal is the difference between the background conductive heat flow and the well heat flow. The reported convective signals contain outliers that may affect upflow prediction, so the influence of outliers is tested by constructing models for two cases: 1) using all the data (i.e., -91 to 11,105 mW/m2), and 2) truncating the range of labels to include only reported convective signals between -25 and 200 mW/m2. Because hydrothermal systems are sparse, models that predict high convective signal in smaller areas better match the natural frequency of hydrothermal systems. Early results demonstrate that XGBoost outperforms linear regression. For XGBoost using the truncated range of labels, half of the high reported signals are within < 3 % of the highest predictions. For XGBoost using the entire range of labels, half of the high reported signals are in < 13 % of the highest predictions. While this implies that the truncated regression is superior, the all-data model better predicts the locations of power-producing systems (i.e., the operating power plants are in a smaller fraction of the study area given by the highest predictions). Even though the models generally predict greater hydrothermal upflow for higher reported convective signals than for lower reported convective signals, both XGBoost models consistently underpredict the magnitude of higher signals. This behavior is attributed to low resolution/granularity of input features compared with the scale of a hydrothermal upflow zone (a few km or less across). Trouble estimating exact values while still reliably predicting high versus low convective signals suggests that a future strategy such as ranked ordinal regression (e.g., classifying into ordered bins for low, medium, high, and very high convective signal) might fit better models, since doing so reduces problems introduced by outliers while preserving the property of larger versus smaller signals.

Geothermal Resources Council Transactions↗

Don’t Let Negatives Hold You Back: Accounting for Underlying Physics and Natural Distributions of Hydrothermal Systems When Selecting Negative Training Sites Leads to Better Machine Learning Predictions

Selecting negative training sites is an important challenge to resolve when utilizing machine learning (ML) for predicting hydrothermal resource favorability because ideal models would discriminate between hydrothermal systems (positives) and all types of locations without hydrothermal systems (negatives). The Nevada Machine Learning project (NVML) fit an artificial neural network to identify areas favorable for hydrothermal systems by selecting 62 negative sites where the research team had confidence that no hydrothermal resource exists. Herein, we compare the implications of the expert selection of negatives (i.e., the NVML strategy) with a random sample strategy, where it is assumed that areas outside the favorable structural ellipses defined by NVML are negative. Because hydrothermal systems are sparse, it is highly probable that, in the absence of a favorable geological structure, hydrothermal favorability is low. We compare three training strategies: 1) the positive and negative labeled examples from NVML; 2) the positive examples from NVML with randomly selected negatives in equal frequency as NVML; and 3) the positive examples from NVML with randomly selected negatives reflecting the expected natural distribution of hydrothermal systems relative to the total area. We apply these training strategies to the NVML feature data (input data) using two ML algorithms (XGBoost and logistic regression) to create six favorability maps for hydrothermal resources. When accounting for the expected natural distribution of hydrothermal systems, we find that XGBoost performs better than the NVML neural network and its negatives. Model validation was less reliable using F1 scores, a common performance metric, than comparing probability estimates at known positives, likely because of the extreme natural class imbalance and the lack of negatively labeled sites. This work demonstrates that expert selection of negatives for training in NVML likely imparted modeling bias. Accounting for the sparsity of hydrothermal systems and all the types of locations without hydrothermal systems allows us to create better models for predicting hydrothermal resource favorability.

Geothermal Resources Council Transactions↗