Geology ReportsSearch

SEARCH · Geology Reports

Results for “Data Science in Science”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,642 records · Page 15Linked to original sources

The feasibility of using lidar-derived digital elevation models for gravity data reduction

Gravity data require submeter elevation accuracy for data processing, and differential global navigation satellite system (dGNSS) equipment is commonly used to acquire three-dimensional positional data to achieve such accuracy. However, lidar (light detection and ranging) data are commonly used to develop digital elevation models (DEMs) of Earth’s surface. Therefore, using elevations from lidar-derived DEMs for gravity-data acquisition and reduction may improve field efficiency and reduce cost. This study examines the feasibility of using DEMs for gravity-data reduction by comparing dGNSS elevation data from 435 gravity stations in Michigan, Wyoming, and Colorado with their respective DEM elevations. The results show that the average difference between DEM and dGNSS elevations is 13 centimeters (cm) and that 93 percent of those differences are less than 50 cm, even in areas with steep terrain. Because an elevation discrepancy of 50 cm corresponds to an error of roughly 0.1 milligals (mGal) in the simple Bouguer gravity anomaly, the results suggest that lidar-derived DEMs are a viable source for acquiring the elevation data needed to process gravity data, thus improving both the cost and efficiency of data collection for regional surveys where an accuracy of less than 1.0 mGal is desired.

Open-File Report

From critical minerals to food security, the benefits of data collaboration

The volume of data in the public geoscience sphere is rapidly and continually expanding. At Geoscience Australia (GA) we saw an over 500% increase in data points within our relational databases between 2018 and 2024, over the life of the Exploring for the Future (EFTF) program. With the Resourcing Australia’s Prosperity initiative, a continued increase in data quantity will be seen for the next 10 to 35 years. At the same time, a broadening audience for geoscience data is increasing the desire to enhance the diversity of delivery streams. This ranges from data-dense highly technical outputs for geoscience specialists to curated interpretive products for people who are non-geoscientists. Development of these curated outputs has contributed to our awareness of the need for data to be collected and compiled in a way that ensures its reuse, with a focus on quality metadata and data provenance.

Conference Paper

A model uncertainty quantification protocol for evaluating the value of observation data

The history-matching approach to parameter estimation with models enables a powerful offshoot analysis of data worth—using the uncertainty of a model forecast as a metric for the worth of data. Adding observation data will either have no impact on forecast uncertainty or will reduce it. Removing existing data will either have no impact on forecast uncertainty or will increase it. The history-matching framework makes it possible to perform this quantitative analysis leveraging the connections among observations, model parameters, and model forecasts. We show this behavior on a specific groundwater flow model of the Mississippi Alluvial Plain and show where the analysis can be informative for considering the potential design of an observation network based on existing or potential observations.

Scientific Investigations Report

Effect of mineral deposit data on predictions from the three-part approach to quantitative mineral resource assessment—A study of 16 previous U.S. Geological Survey assessments

The three-part approach to quantitative mineral resource assessment requires information about the properties of undiscovered mineral deposits in an assessment area. These properties are unknown, so the properties of discovered mineral deposits of the same mineral deposit type are used instead. In the three-part approach, these discovered mineral deposits come from around the world, and their properties constitute the pooled data for that mineral deposit type. Alternatively, these discovered mineral deposits could come from the assessment area, and their properties constitute the tract data for that mineral deposit type. Tract data may be more representative of the undiscovered mineral deposits in the assessment area than the pooled data. The goal of this study was to determine whether resource predictions using pooled data are equivalent to resource predictions using tract data. To this end, 16 previous U.S Geological Survey assessments were studied. For each assessment, resources were predicted for one undiscovered mineral deposit in the assessment area. One set of predictions used pooled data, and another used tract data. The two sets of predictions were compared with an equivalence test, using the six assessment statistics that are commonly reported for mineral resource assessments. Practical equivalence is the condition that two corresponding assessment statistics are within a factor of 1.5 of one another. For each of 2 assessments, all 6 assessment statistics were practically equivalent. For both assessments, the assessment statistics from the pooled data, relative to the corresponding assessment statistics from the tract data, ranged from 1.30 times smaller to 1.03 times larger. For each of 14 assessments, 1 or more of the 6 assessment statistics were not practically equivalent. The assessment statistics from the pooled data, relative to the corresponding assessment statistics from the tract data, ranged from 26.6 times smaller to 5.53 times larger. The use of pooled data has been a standard procedure in the three-part approach since at least 1986. The 16 assessments in this study are not a representative sample of those prior assessments that used pooled data. So, it is inappropriate to use the study results to infer whether pooled data affected the resource predictions for those prior assessments.

Scientific Investigations Report

Detecting earthquakes in noisy real-time GNSS data with deep learning for improved PGD magnitude estimation

To disseminate accurate and useful warnings, earthquake early warning (EEW) systems must quickly determine the size and location of an earthquake to estimate expected shaking. Traditional seismic‐based algorithms tend to underestimate the true magnitudes of large earthquakes, a phenomenon known as magnitude saturation. This limitation motivated the recent inclusion of Global Navigation Satellite Systems (GNSS) data into the U.S. Geological Survey’s ShakeAlert EEW system with the Geodetic First Approximation of Size and Time (GFAST) algorithm because GNSS data do not saturate with large ground motions. However, the noise levels of GNSS data are very high compared with traditional seismic data, which obscures P ‐wave arrivals and can result in less accurate magnitude estimations if displacement amplitudes are low, such as for lower magnitude earthquakes or large source–station distances. In this study, we develop a deep‐learning model that detects earthquakes in GNSS data and use the Ridgecrest, California, earthquake sequence as a case study to demonstrate how the model could act as a filter to reduce the amount of low‐quality data that enters an algorithm like GFAST. To preserve our limited real earthquake data for model inference, we generated a training dataset composed of >700,000 synthetic displacement waveforms. We combined the synthetic waveforms with real‐time GNSS noise to produce realistically noisy training waveforms and then tested our model on additional synthetic data and performed inference using the real data that were held back. We discuss the performance of our trained model on both the unseen synthetic data and real inference data. Our model can be used to selectively filter only high‐quality data where an earthquake signal is observed for input into an algorithm like GFAST (outperforming a simple signal‐to‐noise ratio–based filter) to reduce the error in GFAST’s real‐time earthquake magnitude estimations.

California

A multi-channel digital telemetry system for low frequency geophysical data

An inexpensive general purpose digital telemetry system for collection of low frequency geophysical data from U.S. Geological Survey instruments (eg. tilt, strain, gravity, creep, water level, radon, magnetic field, resistivity, telluric current, temperature, etc.) has been designed and built. This system provides data for a more general interactive data acquisition, retrieval and analysis system. The field stations are self-contained, battery operated and housed in weather proof containers. Each accepts up to 15 analog data inputs in the range of -5 to +5 volts. The dynamic range is 70db. The units transmit information as FSK (Frequency Shift Keyed) tones onto either a phone line or radio link with up to 150 transmitters sharing one line. The average power consumption is 0.06 nR watts where n is the 1 number of input channels transmitted and R is the sample rate in minutes -1 . The central receiver-recorder unit accepts and decodes the FSK tones and converts, formats and records the digital data together with time information and station identification on IBM combatible magnetic tape. The digital data are also converted and recorded in analog form for visual monitoring.

Open-File Report

Relationship of basin structure and bedrock lithology to faulting in the 2019 Ridgecrest earthquake region, California, from gravity and aeromagnetic data

We investigate patterns of cumulative offsets on the faults that ruptured in 2019 and along the Garlock Fault in the Ridgecrest region, California using recently published gravity and aeromagnetic data. We also examine the relationship of basin structure and bedrock structure to the 2019 M7.1 Ridgecrest earthquake ruptures (Fig. 1A), which were primarily along a dextral northwest-striking fault system, and along a sinistral northeast-striking fault, which ruptured hours earlier with a M6.4 event.

California

Toward a new framework to evaluate process-based model configurations and quantify data worth prior to calibration

Model criticism, discrimination, and selection methods often rely on calibrated model outputs. Because calibration can be computationally expensive, model criticism can first be undertaken by assessing model outputs obtained from limited prior parameter ensembles. However, such prior-based methods are often heuristic and do not formalize the notion of balancing model consistency with data and model complexity (i.e., model adequacy). We present a new framework to discriminate among candidate models prior to calibration that formalizes prior-to-calibration model adequacy into a metric to implicitly balance prior model output data coverage with model complexity represented by prior output (co)variance. The prior model adequacy metric “Mahalanobis distance deviation” quantifies the deviation of (a) the set of squared Mahalanobis distances of data from a prior model output distribution from (b) the set of squared Mahalanobis distances of data from their own distribution. A new data worth metric “discernment value” is also presented which quantifies the value of data for screening less-adequate models prior to calibration. Discernment value is calculated from the change in variance of a weighted average of prior model outputs from all candidate models due to less-adequate model outputs receiving lower weight. The framework is demonstrated using a one-dimensional groundwater flow model with eight possible configurations. A synthetic data network is used to test the framework. Results show the framework identifies the candidate models most similar to the true model used to create the synthetic data. Discernment values show variation in the value of different data types and locations for screening less-adequate models.

Water Resources Research

Distinguishing natural sources from anthropogenic events in seismic data

As seismic data are increasingly used to investigate a diverse range of subsurface phenomena beyond regular fast-rupturing earthquakes (Peng and Gomberg, 2010; Beroza and Ide, 2011), it is important to acknowledge that human-generated ground vibrations may be mistaken for naturally generated subsurface processes (Larose et al., 2015; Li et al., 2018). Correct discrimination of natural processes from anthropogenic noise is especially pressing given the trend in seismic detection research toward automated algorithms and machine learning methods (Yoon et al., 2015; Kong et al., 2019;Mousavi and Beroza, 2022) and the growth in seismic data collection in new environments such as urban and industry settings (e.g., Díaz et al.,2017).

Seismological Research Letters

The Sedimentary Geochemistry and Paleoenvironments Project Phase 2 data release: An open data resource for the study of Earth's environmental history

Geochemical data from sedimentary rocks are the primary source of information regarding Earth's surface evolution through time, including its air and water envelopes and interactions with life and deep Earth processes. The Sedimentary Geochemistry and Paleoenvironments Project (SGP) is a scientific consortium centered around open data and community-driven development of cyberinfrastructure tools and resources for sedimentary geochemistry and Earth history. Here we describe the SGP Phase 2 data release, which focused on incorporating Paleoproterozoic and Mesoproterozoic (2500–1000 million years ago) data and better accommodating carbonate data. This data release was built through the involvement of >200 researchers worldwide in academia, government, and industry, and provides the largest available public data resource for our user community in the academic fields of geochemistry, sedimentology, tectonics, paleontology, Earth history, and paleoclimate, as well as the petroleum and minerals industries. The dataset now encompasses 126,006 samples and 4,132,371 geochemical analyses. In addition to direct entry by SGP Team Members, we have ingested and incorporated datasets from the Geoscience Australia OZCHEM database, the Alberta Geological Survey, and the Deep-Time Marine Sedimentary Element Database (DM-SED) compilation. This paper details sampling in the Phase 2 dataset with respect to age, geography, lithology, and other geological characteristics, documents access via our search website and API, discusses possible issues and/or biases in the dataset that could impact analyses, describes plans for governance and stewardship of data from Indigenous lands, and serves as the citable reference paper for the data release.

Chemical Geology

Assessment of density pattern retention of generalized data for 1:100,000-scale United States topographic maps

Cartographic generalization reduces the complexity of geographic data to produce legible, smaller-scale displays that retain essential information and logical geographic patterns. Generalization is a vital process in topographic map production. An important challenge in this process is managing and evaluating consistency across scale in the density and spatial distribution of map features such as buildings, roads, streams, water bodies, and elevation contours. Density patterns in these features reflect underlying physiographic conditions, which include factors such as bedrock geology, tectonics, climate, and landforms. Assessments of an acceptable level of change in feature density patterns are critical to ensuring the readability, usability, and accuracy of generalized maps and data. Preserving realistic density patterns across mapping scales also supports sustainable development goals in cartography, by helping to prioritize and communicate the relative reliability of geospatial data at specific scales.

Conference Paper

New developments at the Center for Engineering Strong-Motion Data (CESMD)

The Center for Engineering Strong-Motion Data (CESMD), an internationally utilized joint center of the U.S. Geological Survey (USGS) and the California Geological Survey (CGS), provides a single access point for earthquake strong-motion records and station metadata from the CGS California Strong-Motion Instrumentation Program (CSMIP), the USGS National Strong-Motion Project (NSMP), the USGS Advanced National Seismic System, and other affiliates. The CESMD has been continuously improving its webtools to facilitate the access of strong-motion data and metadata for use in post-earthquake response and for scientific and engineering research applications. The Center provides raw and processed strong-motion data via the Engineering Data Center (EDC) and the Virtual Data Center (VDC) web portals. This paper focuses on the strong-motion products provided by the EDC where more than 48,000 records with peak ground accelerations greater than 0.1% g from over 2400 earthquakes are currently hosted. and on the ongoing efforts to develop data access tools and applications. The new developments and ongoing efforts in the EDC include: 1) enhancements to the CESMD webservices to facilitate access to station metadata, earthquake information, and strong motion records 2) new features to the interactive map interface, improving the visualization and access to earthquake, station, and record information, 3) efforts to develop a new web application tool for data format conversion from a number of data formats, 4) efforts to unify varying waveform data formats into a consistent format, 5) ongoing efforts to compile seismic station site geology, measured or inferred Vs30 values, shear-wave profiles, NEHRP site class, and available structural instrument deployment schematics, and 6) a special studies pages for research topic-specific ground motion datasets that offer uniform processing of records from a variety of sources.

Conference Paper

U.S. Geological Survey geomagnetic variometer data: Capitalizing on seismic infrastructure

The U.S. Geological Survey’s Geomagnetism Program is collaborating with the Earthquake Hazards Program and Global Seismographic Network Program to densify magnetic field observations. This collaboration focuses on the installation of magnetometers, or magnetic variometers, at existing seismic stations. Along with improving the density of space weather observations for hazard monitoring, these data can be used to correct colocated magnetic field induced noise in seismic data. Such corrections are especially useful during time periods of large magnetic storms where the magnetic field‐induced instrument noise can be of similar amplitude to earthquake ground‐motion records.

contiguous United States

lasertram: A Python library for time resolved analysis of laser ablation inductively coupled plasma mass spectrometry data

Laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) data has a wide variety of uses in the geosciences for in-situ chemical analysis of complex natural materials. Improvements to instrument capabilities and operating software have drastically reduced the time required to generate large volumes of data relative to previous methodologies. Raw data from LA-ICP-MS, however, is in counts per unit time (typically counts per second), not elemental concentrations and converting these count ratesto concentrations requires additional processing. For complex materials where the ablated volume may contain a range of material compositions, a moderate amount of user input is also required if appropriate concentrations are to be accurately calculated. In geologic materials such as glasses and minerals that potentially have numerous heterogeneities (e.g., microlites or other inclusions) within them, this is typically determiningwhether the total ablation signal should be filtered to remove these heterogeneities. This necessitates that the LA-ICP-MS data processing pipeline is one that is not automated, but is also designed to enable rapid and efficient processing of large volumes of data. Here we introduce , a Python library for the time resolved analysis of LA-ICP-MS data. We outline its mathematical theory, code structure, and provide an example of how it can be used to provide the time resolved analysis necessitated by LA-ICP-MS data of complex geologic materials. Throughout the pipeline we show how metadata and data are incrementally added to the objects created such that virtually any aspect of an experiment may be interrogated and its quality assessed. We also show, that when combined with other Python libraries for building graphical user interfaces, it can be utilized outside of a pure scripting environment. can be found at https://doi.org/10.5066/P1DZUR3Z

Applied Computing and Geosciences

Geospatial PDF map of the compilation of GIS data for the mineral industries and related infrastructure of Africa

Introduction In 2021, the U.S. Geological Survey's (USGS) National Minerals Information Center (NMIC) completed the project titled "Compilation of geospatial data for the mineral industries and related infrastructure of Africa." This project aimed to leverage the expertise and capabilities of the NMIC to collect, synthesize, and interpret geospatial data to inform on the extractive resources of the African region and expand the NMIC's understanding on the impact of mineral industry of African nations in the global economy. The African region, which comprises the independent nations that make up the African continent and its associated islands and dependencies, consists of a total of 58 mineral producing countries. The primary objective of this effort was to create a fully attributed Geographic Information System (GIS) portraying existing mining infrastructure, resources, and development capacity across Africa along with the related infrastructure capable of supporting current (for the reference year 2018) and future extractive industry operations in the region. The compiled GIS geodatabase with supporting documentation including comprehensive metadata was published as a USGS data release titled "Compilation of Geospatial Data (GIS) for the Mineral Industries and Related Infrastructure of Africa." This georeferenced portable document format (GeoPDF) map sheet presents a new geographic information product containing a partial representation of the GIS data. This GeoPDF map provides a visual comparison of the distribution of mineral industry and related infrastructure GIS data, which contributes to a deeper understanding of the intersections and complexities of the extractive industries within Africa.

Open-File Report