Geology ReportsSearch

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Assessing the value and usage of data management planning and data management plans within the U.S. Geological Survey

As of 2016, the U.S. Geological Survey (USGS) Fundamental Science Practices require data management plans (DMPs) for all USGS and USGS-funded research. The USGS Science Data Management Branch of the Science Analytics and Synthesis Program has been working to help the USGS (Bureau) meet this requirement. However, USGS researchers still encounter common data management-related challenges that may be reduced or eliminated by better planning. In 2021, USGS staff were given a series of surveys aimed to better understand current data management planning practices, perceptions, and needs. The survey results indicated that adoption and integration of data management planning and DMPs into USGS research project workflows are broad, if inconsistent, across USGS Science Centers and programs. The USGS Science Data Management Branch can help improve clarity and guidance on the purpose, intended audience, content, workflows, and evaluation processes for DMPs. It would also be beneficial to provide additional supporting cyberinfrastructure to support DMP activities. Survey responses indicated it would be beneficial for the Science Data Management Branch to develop a strategy, other than through DMPs, for teaching and encouraging good data management practices. Although these surveys were an opportunity for USGS staff to provide feedback on their experiences, the surveys may also have revealed the desire for more frequent evaluations, cross-disciplinary communication, and training on research data management and DMP development and integration, in the context of USGS policy, Fundamental Science Practices requirements, and overall Bureau expectations. Data management-related roles such as data manager or steward, information technologist, and repository manager may need to be formally recognized as skilled professional career positions within the Bureau. At a minimum, the best practice for USGS would be to create and maintain DMPs as living documents, integrated with existing systems that are broadly accessible to all stakeholders, and include quantitatively measurable benefits tied directly to a clearly defined purpose.

Open-File Report

Definitions of components of the Master Water Data Index maintained by the National Water Data Exchange

The Master Water Data Index (MWDI) is a computerized data base developed and maintained by the National Water Data Exchange (NAWDEX). The MWDI contains information about water-data collection sites. This information includes: the identification of sites for which water data are available, the location of these sites, the type of site, the data-collection organization, the types of data available, the major water-data parameters for which data are available, the frequency at which these parameters are measured, and the media in which the data are stored. This document contains a definition and description of each component of the MWDI data base. (Woodard-USGS)

Open-File Report

Satellite and earth science data management activities at the U.S. geological survey's EROS data center

The U.S. Geological Survey's Earth Resources Observation Systems (EROS) Data Center, the national archive for Landsat data, has 20 years of experience in acquiring, archiving, processing, and distributing Landsat and earth science data. The Center is expanding its satellite and earth science data management activities to support the U.S. Global Change Research Program and the National Aeronautics and Space Administration (NASA) Earth Observing System Program. The Center's current and future data management activities focus on land data and include: satellite and earth science data set acquisition, development and archiving; data set preservation, maintenance and conversion to more durable and accessible archive medium; development of an advanced Land Data Information System; development of enhanced data packaging and distribution mechanisms; and data processing, reprocessing, and product generation systems.

Conference Paper

International data gaps at the Center for Engineering Strong Motion Data

The Center for Engineering Strong Motion Data (CESMD) is utilized by seismologists, engineers, and disaster management professionals in the US and has historically achieved and distributed waveforms from across the globe for significant earthquakes. The increased access to the waveforms via Web API (Application Programming Interface) offers a unique opportunity to provide the community complete datasets, sampling a variety of tectonic environments and geologic conditions, increasing the number of available ground motion records for use in ground motion models (GMMs) and improving the accuracy of earthquake engineering evaluations. The objective of this study is to programmatically identify gaps in global event data from the past decade and backfill missing data gaps at CESMD. We first compare the CESMD catalog with the Advanced National Seismic System (ANSS) Comprehensive Earthquake Catalog identifying regions and time periods where strong-motion data is limited or inadequate. To backfill datasets at CESMD for significant events, we pinpoint regions and time intervals that lack information, creating a list of events for which we’d like to obtain data. An important facet of this work is identifying the source of data and metadata across earthquake repositories around the world and integrating these data repositories into our current strong-motion data processing workflow. In parallel with these newly processed datasets, we are developing a script to produce data origination citations to include provenance and attribution information to associate with respective datasets at CESMD. We showcase our methodology for identifying and filling data gaps at CESMD using three case studies (the 2018 Anchorage Alaska earthquake sequence, seismicity associated with the 2018 Hawaiian Kilauea volcano eruption, and several earthquakes in Turkey) and then outline our strategy to apply our data gap backfilling methods on an international scale.

Conference Paper

The data quality analyzer: a quality control program for seismic data

The U.S. Geological Survey's Albuquerque Seismological Laboratory (ASL) has several initiatives underway to enhance and track the quality of data produced from ASL seismic stations and to improve communication about data problems to the user community. The Data Quality Analyzer (DQA) is one such development and is designed to characterize seismic station data quality in a quantitative and automated manner. The DQA consists of a metric calculator, a PostgreSQL database, and a Web interface: The metric calculator, SEEDscan, is a Java application that reads and processes miniSEED data and generates metrics based on a configuration file. SEEDscan compares hashes of metadata and data to detect changes in either and performs subsequent recalculations as needed. This ensures that the metric values are up to date and accurate. SEEDscan can be run as a scheduled task or on demand. The PostgreSQL database acts as a central hub where metric values and limited station descriptions are stored at the channel level with one-day granularity. The Web interface dynamically loads station data from the database and allows the user to make requests for time periods of interest, review specific networks and stations, plot metrics as a function of time, and adjust the contribution of various metrics to the overall quality grade of the station. The quantification of data quality is based on the evaluation of various metrics (e.g., timing quality, daily noise levels relative to long-term noise models, and comparisons between broadband data and event synthetics). Users may select which metrics contribute to the assessment and those metrics are aggregated into a “grade” for each station. The DQA is being actively used for station diagnostics and evaluation based on the completed metrics (availability, gap count, timing quality, deviation from a global noise model, deviation from a station noise model, coherence between co-located sensors, and comparison between broadband data and synthetics for earthquakes) on stations in the Global Seismographic Network and Advanced National Seismic System.

Computers & Geosciences

What is “big data” and how should we use it? The role of large datasets, secondary data, and associated analysis techniques in outdoor recreation research

With researchers increasingly interested in big data research, this conceptual paper describes how large datasets, secondary data, and associated analysis techniques can be used to understand outdoor recreation. Some types of large, secondary datasets that have been increasingly used in outdoor recreation research include social media, mobile device data, and trip reports or online reviews. First, we give a brief overview of big data terms and outline the steps involved in conducting big data research. In doing so, we describe data sources and analysis techniques relevant for outdoor recreation, and review how they have been applied in previous published works. We then describe opportunities, limitations, and considerations of using big data. Finally, we outline several questions researchers may consider when designing, conducting, reporting, and reviewing outdoor recreation research using big data. Overall, big data approaches can expand our understanding of outdoor recreation and, by addressing key questions, may help researchers harness the strengths of big data while ensuring quality and integrity.

Journal of Outdoor Recreation and Tourism

A trans-dimensional Bayesian Markov chain Monte Carlo algorithm for model assessment using frequency-domain electromagnetic data

A meaningful interpretation of geophysical measurements requires an assessment of the space of models that are consistent with the data, rather than just a single, ‘best’ model which does not convey information about parameter uncertainty. For this purpose, a trans-dimensional Bayesian Markov chain Monte Carlo (MCMC) algorithm is developed for assessing frequencydomain electromagnetic (FDEM) data acquired from airborne or ground-based systems. By sampling the distribution of models that are consistent with measured data and any prior knowledge, valuable inferences can be made about parameter values such as the likely depth to an interface, the distribution of possible resistivity values as a function of depth and non-unique relationships between parameters. The trans-dimensional aspect of the algorithm allows the number of layers to be a free parameter that is controlled by the data, where models with fewer layers are inherently favoured, which provides a natural measure of parsimony and a significant degree of flexibility in parametrization. The MCMC algorithm is used with synthetic examples to illustrate how the distribution of acceptable models is affected by the choice of prior information, the system geometry and configuration and the uncertainty in the measured system elevation. An airborne FDEM data set that was acquired for the purpose of hydrogeological characterization is also studied. The results compare favorably with traditional least-squares analysis, borehole resistivity and lithology logs from the site, and also provide new information about parameter uncertainty necessary for model assessment.

Geophysical Journal International

Building a multi-scaled geospatial temporal ecology database from disparate data sources: Fostering open science through data reuse

Although there are considerable site-based data for individual or groups of ecosystems, these datasets are widely scattered, have different data formats and conventions, and often have limited accessibility. At the broader scale, national datasets exist for a large number of geospatial features of land, water, and air that are needed to fully understand variation among these ecosystems. However, such datasets originate from different sources and have different spatial and temporal resolutions. By taking an open-science perspective and by combining site-based ecosystem datasets and national geospatial datasets, science gains the ability to ask important research questions related to grand environmental challenges that operate at broad scales. Documentation of such complicated database integration efforts, through peer-reviewed papers, is recommended to foster reproducibility and future use of the integrated database. Here, we describe the major steps, challenges, and considerations in building an integrated database of lake ecosystems, called LAGOS (LAke multi-scaled GeOSpatial and temporal database), that was developed at the sub-continental study extent of 17 US states (1,800,000 km 2 ). LAGOS includes two modules: LAGOS GEO , with geospatial data on every lake with surface area larger than 4 ha in the study extent (~50,000 lakes), including climate, atmospheric deposition, land use/cover, hydrology, geology, and topography measured across a range of spatial and temporal extents; and LAGOS LIMNO , with lake water quality data compiled from ~100 individual datasets for a subset of lakes in the study extent (~10,000 lakes). Procedures for the integration of datasets included: creating a flexible database design; authoring and integrating metadata; documenting data provenance; quantifying spatial measures of geographic data; quality-controlling integrated and derived data; and extensively documenting the database. Our procedures make a large, complex, and integrated database reproducible and extensible, allowing users to ask new research questions with the existing database or through the addition of new data. The largest challenge of this task was the heterogeneity of the data, formats, and metadata. Many steps of data integration need manual input from experts in diverse fields, requiring close collaboration.

Connecticut, Delaware, Illinois, Indiana, Iowa, Ma

Data model and relational database design for the New Jersey Water-Transfer Data System (NJWaTr)

The New Jersey Water-Transfer Data System (NJWaTr) is a database design for the storage and retrieval of water-use data. NJWaTr can manage data encompassing many facets of water use, including (1) the tracking of various types of water-use activities (withdrawals, returns, transfers, distributions, consumptive-use, wastewater collection, and treatment); (2) the storage of descriptions, classifications and locations of places and organizations involved in water-use activities; (3) the storage of details about measured or estimated volumes of water associated with water-use activities; and (4) the storage of information about data sources and water resources associated with water use. In NJWaTr, each water transfer occurs unidirectionally between two site objects, and the sites and conveyances form a water network. The core entities in the NJWaTr model are site, conveyance, transfer/volume, location, and owner. Other important entities include water resource (used for withdrawals and returns), data source, permit, and alias. Multiple water-exchange estimates based on different methods or data sources can be stored for individual transfers. Storage of user-defined details is accommodated for several of the main entities. Many tables contain classification terms to facilitate the detailed description of data items and can be used for routine or custom data summarization. NJWaTr accommodates single-user and aggregate-user water-use data, can be used for large or small water-network projects, and is available as a stand-alone Microsoft? Access database. Data stored in the NJWaTr structure can be retrieved in user-defined combinations to serve visualization and analytical applications. Users can customize and extend the database, link it to other databases, or implement the design in other relational database applications.

Open-File Report

Description and testing of the Geo Data Portal: Data integration framework and Web processing services for environmental science collaboration

Interest in sharing interdisciplinary environmental modeling results and related data is increasing among scientists. The U.S. Geological Survey Geo Data Portal project enables data sharing by assembling open-standard Web services into an integrated data retrieval and analysis Web application design methodology that streamlines time-consuming and resource-intensive data management tasks. Data-serving Web services allow Web-based processing services to access Internet-available data sources. The Web processing services developed for the project create commonly needed derivatives of data in numerous formats. Coordinate reference system manipulation and spatial statistics calculation components implemented for the Web processing services were confirmed using ArcGIS 9.3.1, a geographic information science software package. Outcomes of the Geo Data Portal project support the rapid development of user interfaces for accessing and manipulating environmental data.

Open-File Report

Sharing our data—An overview of current (2016) USGS policies and practices for publishing data on ScienceBase and an example interactive mapping application

This report provides an overview of current (2016) U.S. Geological Survey policies and practices related to publishing data on ScienceBase, and an example interactive mapping application to display those data. ScienceBase is an integrated data sharing platform managed by the U.S. Geological Survey. This report describes resources that U.S. Geological Survey Scientists can use for writing data management plans, formatting data, and creating metadata, as well as for data and metadata review, uploading data and metadata to ScienceBase, and sharing metadata through the U.S. Geological Survey Science Data Catalog. Because data publishing policies and practices are evolving, scientists should consult the resources cited in this paper for definitive policy information. An example is provided where, using the content of a published ScienceBase data release that is associated with an interpretive product, a simple user interface is constructed to demonstrate how the open source capabilities of the R programming language and environment can interact with the properties and objects of the ScienceBase item and be used to generate interactive maps.

Open-File Report

Instructions for the submission of data to the Master Water Data Index

Instructions are given for the computerized processing of data prepared for entry into the Master Water Data Index of the U.S. Geological Survey 's National Water Data Exchange (NAWDEX), as discussed in the manual entitled ' Instructions for the Preparation of Data for the Master Water Data Index. ' Two processing procedures are provided: The edit of data and the submission of data for storage in the Master Water Data Index. These instructions are applicable to both new data being submitted for storage and to update transactions for data previously processed and stored. (USGS)

Open-File Report

Metadata for data rescue and data at risk

Scientific data age, become stale, fall into disuse and run tremendous risks of being forgotten and lost. These problems can be addressed by archiving and managing scientific data over time, and establishing practices that facilitate data discovery and reuse. Metadata documentation is integral to this work and essential for measuring and assessing high priority data preservation cases. The International Council for Science: Committee on Data for Science and Technology (CODATA) has a newly appointed Data-at-Risk Task Group (DARTG), participating in the general arena of rescuing data. The DARTG primary objective is building an inventory of scientific data that are at risk of being lost forever. As part of this effort, the DARTG is testing an approach for documenting endangered datasets. The DARTG is developing a minimal and easy to use set of metadata properties for sufficiently describing endangered data, which will aid global data rescue missions. The DARTG metadata framework supports rapid capture, and easy documentation, across an array of scientific domains. This paper reports on the goals and principles supporting the DARTG metadata schema, and provides a description of the preliminary implementation.

Conference Paper

Availability of Earth observations data from the U.S. Geological Survey's EROS data center

For decades federal and state agencies have been collecting regional, continental, and global Earth observations data acquired by satellites, aircraft, and other information-gathering systems. These data include photographic and digital remotely sensed images of the Earth's surface, as well as earth science, cartographic, and geographic data. Since 1973, the U.S. Geological Survey's Earth Resources Observation Systems (EROS) Data Center (EDC) in Sioux Falls, South Dakota, has been a data management, production, dissemination, and research center for these data. Currently, the Data Center holds over 10 million satellite images and aerial photographs, in photographic and digital formats. Users are able to place inquiries and orders for these holdings via a nationwide computer network. In addition to cataloging the data stored in its archives, the Data Center provides users with rapid access to information on many data collections held by other facilities.

Pecora 12 Symposium

The Sedimentary Geochemistry and Paleoenvironments Project Phase 2 data release: An open data resource for the study of Earth's environmental history

Geochemical data from sedimentary rocks are the primary source of information regarding Earth's surface evolution through time, including its air and water envelopes and interactions with life and deep Earth processes. The Sedimentary Geochemistry and Paleoenvironments Project (SGP) is a scientific consortium centered around open data and community-driven development of cyberinfrastructure tools and resources for sedimentary geochemistry and Earth history. Here we describe the SGP Phase 2 data release, which focused on incorporating Paleoproterozoic and Mesoproterozoic (2500–1000 million years ago) data and better accommodating carbonate data. This data release was built through the involvement of >200 researchers worldwide in academia, government, and industry, and provides the largest available public data resource for our user community in the academic fields of geochemistry, sedimentology, tectonics, paleontology, Earth history, and paleoclimate, as well as the petroleum and minerals industries. The dataset now encompasses 126,006 samples and 4,132,371 geochemical analyses. In addition to direct entry by SGP Team Members, we have ingested and incorporated datasets from the Geoscience Australia OZCHEM database, the Alberta Geological Survey, and the Deep-Time Marine Sedimentary Element Database (DM-SED) compilation. This paper details sampling in the Phase 2 dataset with respect to age, geography, lithology, and other geological characteristics, documents access via our search website and API, discusses possible issues and/or biases in the dataset that could impact analyses, describes plans for governance and stewardship of data from Indigenous lands, and serves as the citable reference paper for the data release.

Chemical Geology

Landsat Data Continuity Mission (LDCM) space to ground mission data architecture

The Landsat Data Continuity Mission (LDCM) is a scientific endeavor to extend the longest continuous multi-spectral imaging record of Earth's land surface. The observatory consists of a spacecraft bus integrated with two imaging instruments; the Operational Land Imager (OLI), built by Ball Aerospace & Technologies Corporation in Boulder, Colorado, and the Thermal Infrared Sensor (TIRS), an in-house instrument built at the Goddard Space Flight Center (GSFC). Both instruments are integrated aboard a fine-pointing, fully redundant, spacecraft bus built by Orbital Sciences Corporation, Gilbert, Arizona. The mission is scheduled for launch in January 2013. This paper will describe the innovative end-to-end approach for efficiently managing high volumes of simultaneous realtime and playback of image and ancillary data from the instruments to the reception at the United States Geological Survey's (USGS) Landsat Ground Network (LGN) and International Cooperator (IC) ground stations. The core enabling capability lies within the spacecraft Command and Data Handling (C&DH) system and Radio Frequency (RF) communications system implementation. Each of these systems uniquely contribute to the efficient processing of high speed image data (up to 265Mbps) from each instrument, and provide virtually error free data delivery to the ground. Onboard methods include a combination of lossless data compression, Consultative Committee for Space Data Systems (CCSDS) data formatting, a file-based/managed Solid State Recorder (SSR), and Low Density Parity Check (LDPC) forward error correction. The 440 Mbps wideband X-Band downlink uses Class 1 CCSDS File Delivery Protocol (CFDP), and an earth coverage antenna to deliver an average of 400 scenes per day to a combination of LGN and IC ground stations. This paper will also describe the integrated capabilities and processes at the LGN ground stations for data reception using adaptive filtering, and the mission operations approach fro- the LDCM Mission Operations Center (MOC) to perform the CFDP accounting, file retransmissions, and management of the autonomous features of the SSR.

Conference Paper

Description of the U.S. Geological Survey Geo Data Portal data integration framework

The U.S. Geological Survey has developed an open-standard data integration framework for working efficiently and effectively with large collections of climate and other geoscience data. A web interface accesses catalog datasets to find data services. Data resources can then be rendered for mapping and dataset metadata are derived directly from these web services. Algorithm configuration and information needed to retrieve data for processing are passed to a server where all large-volume data access and manipulation takes place. The data integration strategy described here was implemented by leveraging existing free and open source software. Details of the software used are omitted; rather, emphasis is placed on how open-standard web services and data encodings can be used in an architecture that integrates common geographic and atmospheric data.

IEEE Journal of Selected Topics in Applied Earth O