Geology ReportsSearch

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Database of well and areal data, South San Francisco Bay and Peninsula area, California

A database was developed to organize and manage data compiled for a regional assessment of geohydrologic and water-quality conditions in the south San Francisco Bay and Peninsula area in California. Available data provided by local, State, and Federal agencies and private consultants was utilized in the assessment. The database consists of geographicinformation system data layers and related tables and American Standard Code for Information Interchange files. Documentation of the database is necessary to avoid misinterpretation of the data and to make users aware of potential errors and limitations. Most of the data compiled were collected from wells and boreholes (collectively referred to as wells in this report). This point-specific data, including construction, water-level, waterquality, pumping test, and lithologic data, are contained in tables and files that are related to a geographic information system data layer that contains the locations of the wells. There are 1,014 wells in the data layer and the related tables contain 35,845 water-level measurements (from 293 of the wells) and 9,292 water-quality samples (from 394 of the wells). Calculation of hydraulic heads and gradients from the water levels can be affected adversely by errors in the determination of the altitude of land surface at the well. Cation and anion balance computations performed on 396 of the water-quality samples indicate high cation and anion balance errors for 51 (13 percent) of the samples. Well drillers' reports were interpreted for 762 of the wells, and digital representations of the lithology of the formations are contained in files following the American Standard Code for Information Interchange. The usefulness of drillers' descriptions of the formation lithology is affected by the detail and thoroughness of the drillers' descriptions, as well as the knowledge, experience, and vocabulary of the individual who described the drill cuttings. Additional data layers were created that contain political, geohydrologic, and other geographic data. These layers contain features represented by areas and lines rather than discrete points. The layers consist of data representing the thickness of alluvium, surficial geology, physiographic subareas, watershed boundaries, land use, water-supply districts, wastewater treatment districts, and recharge basins. The layers manually digitizing paper maps, acquisition of data already in digital form, or creation of new layers from available layers. The scale of the source data affects the accurate representation of real-world features with the data layer, and, therefore, the scale of the source data must be considered when the data are analyzed and plotted.

Water-Resources Investigations Report

Seven recommendations to make your invasive alien species data more useful

Science-based strategies to tackle biological invasions depend on recent, accurate, well-documented, standardized and openly accessible information on alien species. Currently and historically, biodiversity data are scattered in numerous disconnected data silos that lack interoperability. The situation is no different for alien species data, and this obstructs efficient retrieval, combination, and use of these kinds of information for research and policy-making. Standardization and interoperability are particularly important as many alien species related research and policy activities require pooling data. We describe seven ways that data on alien species can be made more accessible and useful, based on the results of a European Cooperation in Science and Technology (COST) workshop: (1) Create data management plans; (2) Increase interoperability of information sources; (3) Document data through metadata; (4) Format data using existing standards; (5) Adopt controlled vocabularies; (6) Increase data availability; and (7) Ensure long-term data preservation. We identify four properties specific and integral to alien species data (species status, introduction pathway, degree of establishment, and impact mechanism) that are either missing from existing data standards or lack a recommended controlled vocabulary. Improved access to accurate, real-time and historical data will repay the long-term investment in data management infrastructure, by providing more accurate, timely and realistic assessments and analyses. If we improve core biodiversity data standards by developing their relevance to alien species, it will allow the automation of common activities regarding data processing in support of environmental policy. Furthermore, we call for considerable effort to maintain, update, standardize, archive, and aggregate datasets, to ensure proper valorization of alien species data and information before they become obsolete or lost.

Frontiers in Applied Mathematics and Statistics

Alaska Geochemical Database (AGDB)-Geochemical data for rock, sediment, soil, mineral, and concentrate sample media

The Alaska Geochemical Database (AGDB) was created and designed to compile and integrate geochemical data from Alaska in order to facilitate geologic mapping, petrologic studies, mineral resource assessments, definition of geochemical baseline values and statistics, environmental impact assessments, and studies in medical geology. This Microsoft Access database serves as a data archive in support of present and future Alaskan geologic and geochemical projects, and contains data tables describing historical and new quantitative and qualitative geochemical analyses. The analytical results were determined by 85 laboratory and field analytical methods on 264,095 rock, sediment, soil, mineral and heavy-mineral concentrate samples. Most samples were collected by U.S. Geological Survey (USGS) personnel and analyzed in USGS laboratories or, under contracts, in commercial analytical laboratories. These data represent analyses of samples collected as part of various USGS programs and projects from 1962 to 2009. In addition, mineralogical data from 18,138 nonmagnetic heavy mineral concentrate samples are included in this database. The AGDB includes historical geochemical data originally archived in the USGS Rock Analysis Storage System (RASS) database, used from the mid-1960s through the late 1980s and the USGS PLUTO database used from the mid-1970s through the mid-1990s. All of these data are currently maintained in the Oracle-based National Geochemical Database (NGDB). Retrievals from the NGDB were used to generate most of the AGDB data set. These data were checked for accuracy regarding sample location, sample media type, and analytical methods used. This arduous process of reviewing, verifying and, where necessary, editing all USGS geochemical data resulted in a significantly improved Alaska geochemical dataset. USGS data that were not previously in the NGDB because the data predate the earliest USGS geochemical databases, or were once excluded for programmatic reasons, are included here in the AGDB and will be added to the NGDB. The AGDB data provided here are the most accurate and complete to date, and should be useful for a wide variety of geochemical studies. The AGDB data provided in the linked database may be updated or changed periodically. The data on the DVD and in the data downloads provided with this report are current as of date of publication.

Data Series

Data resources for range-wide assessment of livestock grazing across the sagebrush biome

The data contained in this series were compiled, modified, and analyzed for the U.S. Geological Survey (USGS) report "Range-Wide Assessment of Livestock Grazing Across the Sagebrush Biome." This report can be accessed through the USGS Publications Warehouse (online linkage: http://pubs.usgs.gov/of/2011/1263/). The dataset contains spatial and tabular data related to Bureau of Land Management (BLM) Grazing Allotments. We reviewed the BLM national grazing allotment spatial dataset available from the GeoCommunicator National Integrated Land System (NILS) website in 2007 (http://www.geocommunicator.gov). We identified several limitations in those data and learned that some BLM State and/or field offices had updated their spatial data to rectify these limitations, but maintained the data outside of NILS. We contacted appropriate BLM offices (State or field, 25 in all) to obtain the most recent data, assessed the data, established a data development protocol, and compiled data into a topologically enforced dataset throughout the area of interest for this project (that is, the pre-settlement distribution of Greater Sage-Grouse in the Western United States). The final database includes three spatial datasets: Allotments (BLM Grazing Allotments), OUT_Polygons (nonallotment polygons used to ensure topology), and Duplicate_Polygon_Allotments. See Appendix 1 of the aforementioned report for complete methods. The tabular data presented here consists of information synthesized by the Land Health Standard (LHS) analysis (Appendix 2), and data obtained from the BLM Rangeland Administration System (http://www.blm.gov/ras/). In 2008, available LHS data for all allotments in all regions were compiled by BLM in response to a Freedom of Information Act (FOIA) request made by a private organization. The BLM provided us with a copy of these data. These data provided three major types of information that were of interest: (1) date(s) (if any) of the most recent LHS evaluation for each allotment; (2) whether if evaluated, each region-specific standard (3–8 LHS depending on region) had been met on a given allotment; and (3) whether livestock contributed to any of these standards not being met. A description of how we processed the original data to prepare for analysis is described in Appendix 2, and the synthesized dataset can be found in the table "lhs_x_walk." Permitted use dates, livestock type (horse, sheep or cattle), number of livestock, and Animal Unit Months [the number of animal units (1,000-pound animal equivalents) that can be grazed for 31 days with the available forage in a sustainable manner] are the legal maximum grazing amounts for a given allotment, and legal adjustments to these numbers occur infrequently. We summarized permitted use by BLM allotment in the table "Permitted_Use." Billed use records are used for calculations of permittees' annual grazing bills. We summarized billed use by allotment for BLM grazing year in the table "Billed_Use." All three tables can be joined with the allotment spatial data in a geographic information system (GIS) environment, using the IDENT attribute as the primary key.

Data Series

Evaluation of downscaled, gridded climate data for the conterminous United States

Weather and climate affect many ecological processes, making spatially continuous yet fine-resolution weather data desirable for ecological research and predictions. Numerous downscaled weather data sets exist, but little attempt has been made to evaluate them systematically. Here we address this shortcoming by focusing on four major questions: (1) How accurate are downscaled, gridded climate data sets in terms of temperature and precipitation estimates?, (2) Are there significant regional differences in accuracy among data sets?, (3) How accurate are their mean values compared with extremes?, and (4) Does their accuracy depend on spatial resolution? We compared eight widely used downscaled data sets that provide gridded daily weather data for recent decades across the United States. We found considerable differences among data sets and between downscaled and weather station data. Temperature is represented more accurately than precipitation, and climate averages are more accurate than weather extremes. The data set exhibiting the best agreement with station data varies among ecoregions. Surprisingly, the accuracy of the data sets does not depend on spatial resolution. Although some inherent differences among data sets and weather station data are to be expected, our findings highlight how much different interpolation methods affect downscaled weather data, even for local comparisons with nearby weather stations located inside a grid cell. More broadly, our results highlight the need for careful consideration among different available data sets in terms of which variables they describe best, where they perform best, and their resolution, when selecting a downscaled weather data set for a given ecological application.

Ecological Applications

What is the (real) rate of soil health practice adoption? Making sense of three data sources

Conservation stakeholders looking to quantify the impact of their investments to increase soil health practice adoption over time often face challenges in interpreting practice adoption data due to discrepancies in language and results among data sources. Similarly, efforts to estimate environmental outcomes of practice adoption, such as water quality and greenhouse gas emissions, can vary depending on different practice adoption input data. To help make sense of different adoption data sources, we compared county-level adoption data for winter cover crops (WCC), no-till (NT), and reduced tillage (RT) in three areas of the United States with contrasting climates and production systems: central Illinois (CIL), southern Illinois (SIL), and western New York (WNY). We analyzed data available during 2015 through 2022 from the Operational Tillage Information System (OpTIS, remote sensing), US Census of Agriculture (AgCensus, a farmer survey), and, specifically in Illinois, the Illinois Soil Conservation Transect Survey (Transect, a roadside survey). The magnitude of differences between the datasets depended on the practice and geographic location. For example, OpTIS and AgCensus tillage data were much more similar in Illinois (average difference of less than 4 percentage points) compared to New York (average differences of 20 percentage points). Similarly, there was less variability and smaller differences between OpTIS and AgCensus WCC data in Illinois compared to WNY. AgCensus tended to report lower WCC adoption for Illinois and greater adoption in WNY compared to OpTIS. All data sources agreed that the rate of change in tillage practices is slow (mainly –1% to 1%) and that adoption of WCC is low (assuming linear growth, it could take nearly a century to reach 50% WCC adoption in CIL). Differences among the datasets were attributed to definitional inconsistencies for RT and NT and how WCC data were acquired. For example, the AgCensus asks if a WCC was planted, whereas OpTIS and Transect evaluate the presence of a standing WCC. Data sources also reflect different time periods (calendar years or crop years) and types of cropland assessed (corn [ Zea mays L.], soybean [ Glycine max {L.} Merr.], or all cropland). We propose two recommendations to improve interpretation and consistency: (1) a working group to harmonize definitions and protocols and develop educational materials for data users, and (2) a research effort that integrates different adoption data types and produces publicly available adoption data at HUC-10 and county scales. Such activities could help improve data access and utility for evidence-based conservation decision-making and enhance the accuracy of environmental models that rely on adoption data as input.

Journal of Soil and Water Conservation

Challenges and opportunities for data integration to improve estimation of migratory connectivity

Understanding migratory connectivity, or the linkage of populations between seasons, is critical for effective conservation and management of migratory wildlife. A growing number of tools are available for understanding where migratory individuals and populations occur throughout the annual cycle. Integration of the diverse measures of migratory movements can help elucidate migratory connectivity patterns with methodology that accounts for differences in sampling design, directionality, effort, precision and bias inherent to each data type. The R package MigConnectivity was developed to estimate population-specific connectivity and the range-wide strength of those connections. New functions allow users to integrate intrinsic markers, tracking and long-distance reencounter data, collected from the same or different individuals, to estimate population-specific transition probabilities (estTransition) and the range-wide strength of those transition probabilities (estStrength). We used simulation and real-world case studies to explore the challenges and limitations of data integration based on data from three migratory bird species, Painted Bunting ( Passerina ciris ), Yellow Warbler ( Setophaga petechia ) and Bald Eagle ( Haliaeetus leucocephalus ), two of which had bidirectional data. We found data integration is useful for quantifying migratory connectivity, as single data sources are less likely to be available across the species range. Furthermore, accurate strength estimates can be obtained from either breeding-to-nonbreeding or nonbreeding-to-breeding data. For bidirectional data, integration can lead to more accurate estimates when data are available from all regions in at least one season. The ability to conduct combined analyses that account for the unique limitations and biases of each data type is a promising possibility for overcoming the challenge of range-wide coverage that has been hard to achieve using single data types. The best-case scenario for data integration is to have data from all regions, especially if the question is range-wide or data are bidirectional. Multiple data types on animal movements are becoming increasingly available and integration of these growing datasets will lead to a better understanding of the full annual cycle of migratory animals.

Methods in Ecology and Evolution

Digital data grids for the magnetic anomaly map of North America

The digital magnetic anomaly database and map for the North American continent is the result of a joint effort by the Geological Survey of Canada (GSC), U. S. Geological Survey (USGS), and Consejo de Recursos Minerales of Mexico (CRM). This integrated, readily accessible, modern digital database of magnetic anomaly data is a powerful tool for further evaluation of the structure, geologic processes, and tectonic evolution of the continent and may also be used to help resolve societal and scientific issues that span national boundaries. The North American magnetic anomaly map derived from the digital database provides a comprehensive magnetic view of continental-scale trends not available in individual data sets, helps link widely separated areas of outcrop, and unifies disparate geologic studies. This open-file report presents three unique, gridded data sets used to make the magnetic anomaly map of North America. Subsets of these three grids that span only the United States were also created, giving a total of six grids. Details on the data processing and compilation procedures used to produce the grids are described in the booklet that accompanies the North American magnetic anomaly map. All three grids have 1-km spacing and are projected to the DNAG projection (spherical transverse mercator, central meridian of 100 o W, base latitude of 0o, scale factor of 0.926 and Earth radius of 6,371,204 m.) More details are given in the metadata files that accompany the gridded data files. These grids are presented in Geosoft binary grid format, with two files describing each of the six grids (suffixes .grd and .gi). This format can be easily converted to numerous other formats using the free conversion software offered by this company at http://www.geosoft.com/. The first grids (NAmag_origmrg.grd and USmag_origmrg.grd) show the magnetic field at 305 m. above terrain. For the second grids (NAmag_hp500.grd and USmag_hp500.grd) we removed long-wavelength anomalies (500 km and greater) from the first grid. This grid was used for the published map. Although the North American merged grid represents a significant upgrade to older compilations, the existing patchwork of surveys is inherently unable to accurately represent anomalies with long (greater than roughly 150 km) wavelengths, particularly in the US and Canada (U.S. Magnetic-Anomaly Data Set Task Group, 1994). The lack of information about long wavelength anomalies is primarily related to datum shifts between merged surveys, caused by data acquisition at widely different times and by differences in merging procedures. Therefore, we removed anomalies with wavelengths greater than 500 km from the merged grid to reduce the effects caused by the spurious long wavelengths but still maintain the continuity of anomalies. The correction was accomplished by transforming the merged grid to the frequency domain, filtering the transformed data with a long-wavelength cutoff at 500 km, and subtracting the long-wavelength data grid from the merged grid. In addition to the 500-km high pass filter, an equivalent source method, based on long-wavelength characterization using satellite data (CHAMP satellite anomalies, Maus and others, 2002), was also used to correct for spurious shifts in the original magnetic anomaly grid (Ravat and others, 2002). These results are presented in the third grids (NAmag_CM.grd and USmag_CM.grd), in which the wavelengths longer than 500 km have been replaced by downward-continued satellite data. The steps used to create the third long-wavelength-corrected grid are: 0. The North American 1-km merged grid was decimated to 5 km. 1. This 5-km grid was converted to a 0.05 degree grid and was low-pass filtered using a Gaussian filter with a 500-km cutoff, then decimated to 1 degree. 2. A joint inversion of this 1-degree low-pass aeromagnetic grid and satellite data, with the aeromagnetic data weighted very low, was used to produce a stabilized downward continuation of the satellite data. 3. The inverted data were interpolated to 0.05 degrees and again low-pass filtered using the same Gaussian 500-km filter to remove short-wavelength artifacts. 4. The low-pass grid from step 1 was subtracted from the original 0.05-degree aeromagnetic grid to create a 500-km high-pass aeromagnetic grid. This grid was added to the low-pass inverted grid from step 3 to get a corrected 0.05-degree aeromagnetic grid. 5. The corrected 0.05-degree aeromagnetic grid was projected to the DNAG projection and regridded to 5 km. This was subtracted from the decimated 5-km aeromagnetic grid to generate a 5-km correction grid. A matched filter was used to remove short-wavelength artifacts resulting from the projection and regridding process. 6. The resulting 5-km correction grid was regridded to the original 1-km grid and subtracted from the original 1-km aeromagnetic grid to generate the final 1-km corrected aeromagnetic grid. The six grids described in this report are available for download. Two metadata files, one for the North American grids and one for the United States grids, are also included with the gridded data.

Open-File Report

Wave data processing toolbox manual

Researchers routinely deploy oceanographic equipment in estuaries, coastal nearshore environments, and shelf settings. These deployments usually include tripod-mounted instruments to measure a suite of physical parameters such as currents, waves, and pressure. Instruments such as the RD Instruments Acoustic Doppler Current Profiler (ADCP(tm)), the Sontek Argonaut, and the Nortek Aquadopp(tm) Profiler (AP) can measure these parameters. The data from these instruments must be processed using proprietary software unique to each instrument to convert measurements to real physical values. These processed files are then available for dissemination and scientific evaluation. For example, the proprietary processing program used to process data from the RD Instruments ADCP for wave information is called WavesMon. Depending on the length of the deployment, WavesMon will typically produce thousands of processed data files. These files are difficult to archive and further analysis of the data becomes cumbersome. More imperative is that these files alone do not include sufficient information pertinent to that deployment (metadata), which could hinder future scientific interpretation. This open-file report describes a toolbox developed to compile, archive, and disseminate the processed wave measurement data from an RD Instruments ADCP, a Sontek Argonaut, or a Nortek AP. This toolbox will be referred to as the Wave Data Processing Toolbox. The Wave Data Processing Toolbox congregates the processed files output from the proprietary software into two NetCDF files: one file contains the statistics of the burst data and the other file contains the raw burst data (additional details described below). One important advantage of this toolbox is that it converts the data into NetCDF format. Data in NetCDF format is easy to disseminate, is portable to any computer platform, and is viewable with public-domain freely-available software. Another important advantage is that a metadata structure is embedded with the data to document pertinent information regarding the deployment and the parameters used to process the data. Using this format ensures that the relevant information about how the data was collected and converted to physical units is maintained with the actual data. EPIC-standard variable names have been utilized where appropriate. These standards, developed by the NOAA Pacific Marine Environmental Laboratory (PMEL) (http://www.pmel.noaa.gov/epic/), provide a universal vernacular allowing researchers to share data without translation.

Open-File Report

Using the U.S. Geological Survey National Water Quality Laboratory LT-MDL to Evaluate and Analyze Data

A long-term method detection level (LT-MDL) and laboratory reporting level (LRL) are used by the U.S. Geological Survey?s National Water Quality Laboratory (NWQL) when reporting results from most chemical analyses of water samples. Changing to this method provided data users with additional information about their data and often resulted in more reported values in the low concentration range. Before this method was implemented, many of these values would have been censored. The use of the LT-MDL and LRL presents some challenges for the data user. Interpreting data in the low concentration range increases the need for adequate quality assurance because even small contamination or recovery problems can be relatively large compared to concentrations near the LT-MDL and LRL. In addition, the definition of the LT-MDL, as well as the inclusion of low values, can result in complex data sets with multiple censoring levels and reported values that are less than a censoring level. Improper interpretation or statistical manipulation of low-range results in these data sets can result in bias and incorrect conclusions. This document is designed to help data users use and interpret data reported with the LTMDL/ LRL method. The calculation and application of the LT-MDL and LRL are described. This document shows how to extract statistical information from the LT-MDL and LRL and how to use that information in USGS investigations, such as assessing the quality of field data, interpreting field data, and planning data collection for new projects. A set of 19 detailed examples are included in this document to help data users think about their data and properly interpret lowrange data without introducing bias. Although this document is not meant to be a comprehensive resource of statistical methods, several useful methods of analyzing censored data are demonstrated, including Regression on Order Statistics and Kaplan-Meier Estimation. These two statistical methods handle complex censored data sets without resorting to substitution, thereby avoiding a common source of bias and inaccuracy.

Open-File Report

Community for Data Integration 2014 annual report

The U.S. Geological Survey (USGS) researches Earth science to help address complex issues affecting society and the environment. In 2006, the USGS held the first Scientific Information Management Workshop to bring together staff from across the organization to discuss the data and information management issues affecting the integration and delivery of Earth science research and investigate the use of “communities of practice” as mechanisms to share expertise about these issues. Out of this effort emerged the Council for Data Integration, which was conceived as an official organizational function that would help guide data integration activities and formalize communities of practice into working groups; however, by 2009 it became evident that many members of the Council for Data Integration had an interest in developing data integration solutions and sharing expertise in a less formal, grassroots manner, which transformed the Council into a Community for Data Integration (CDI). As of 2014, the CDI represents a dynamic community of practice focused on advancing science data and information management and integration capabilities across the USGS and the CDI community. The CDI fosters an environment for collaboration and sharing by bringing together expertise from external partners and representatives across the USGS who are involved in research, data management, and information technology. Membership is voluntary and open to USGS employees and other individuals and organizations willing to contribute to the community (if interested, contact cdi@usgs.gov). The purpose of the CDI is to do the following: • advance understanding of Earth systems through enhanced use of data and information including associated tools and techniques, • provide a forum for people doing work with data integration to come together to share ideas and learn new skills and techniques, and • grow overall USGS capabilities with data and information by increasing visibility of the work of many people throughout the USGS and the CDI community. To achieve these goals, the CDI operates within four applied areas: monthly forums, annual workshop/webinar series, working groups, and projects. The monthly forums, also known as the Opportunity/Challenge of the Month, provide an open dialogue to share and learn about data integration efforts or to present problems that invite the community to offer solutions, advice, and support. Since 2010, the CDI has also sponsored annual workshops/webinar series to encourage the exchange of ideas, sharing of activities, presentations of current projects, and networking among members. Stemming from common interests, the working groups are focused on efforts to address data management and technical challenges including the development of standards and tools, improving interoperability and information infrastructure, and data preservation within USGS and its partners. The growing support for the activities of the working groups led to the CDI’s first formal request for proposals (RFP) process in 2013 to fund projects that produced tangible products. As of 2014, the CDI continues to hold an annual RFP that creates data management tools and practices, collaboration tools, and training in support of data integration and delivery.

Open-File Report

A computerized data base of nitrate concentrations in Indiana ground water

As part of a cooperative study with the Indiana Department of Environmental Management, the U.S. Geological Survey compiled a computerized data base of nitrate concentrations in Indiana ground water. The data included nitrate determinations from more than 29 studies by five Federal and State agencies during June 1973 through August 1991. The National Water Information System software of the U.S. Geological Survey was used to store the data at the U.S. Geological Survey office in Indianapolis, Indiana. Electronic data sets were converted to a standard format of well data, sample data, and analytical data. Data were screened by several error-checking procedures before they were retained in the data base; they were examined for potential duplicates of well location and name. The data base of nitrate concentrations in Indiana ground water contains records of 5,525 samples collected from 4,448 wells in 88 of 92 counties during 1973-91. Those wells included 3,832 drinking-water wells; 536 monitoring wells, 38 livestock-supply wells; and 42 irrigation wells. Nitrate concentrations greater than minimum reporting limits of 0.0 to 0.5 milligrams per liter (mg/L) were determined in 2,453 samples (44 percent of the total). Nitrate in ground water at concentrations greater than 3 mg/L have been considered to be the result of human activities. Nitrate concentrations ranged from 0.005 to 380 mg/L with a median nitrate concentration of 0.3 mg/L. Nitrate concentrations were greater than or equal to 3 mg/L in 704 samples (13 percent of the total). Nitrate concentrations were greater than or equal to the U.S. Environmental Protection Agency Maximum Contaminant Level of 10 mg/L in 188 samples (3.4 percent of the total). Of the 3,832 drinking- water wells in the data base, 147 had at least one sample in which a nitrate concentration was greater than the Maximum Contaminant Level. The percentage of samples with nitrate concentrations greater than or equal to 3 mg/L and greater than or equal to 10 mg/L generally increased during the period 1973 through 1991. The nitrate data base was compiled from numerous data sets that were readily accessible in electronic format. The uses of these data may be limited because they were neither comprehensive nor of a single statistical design. Nonetheless, the nitrate data can be used in several ways: (1) to identify geographic areas with and without nitrate data; (2) to evaluate assumptions, models, and maps of ground-water-contamination potential; and (3) to investigate the relation between environmental factors, land-use types, and the occurrence of nitrate.

Indiana

Global cropland-extent product at 30-m resolution (GCEP30) derived from Landsat satellite time-series data for the year 2015 using multiple machine-learning algorithms on Google Earth Engine cloud

Executive Summary Global food and water security analysis and management require precise and accurate global cropland-extent maps. Existing maps have limitations, in that they are (1) mapped using coarse-resolution remote-sensing data, resulting in the lack of precise mapping location of croplands and their accuracies; (2) derived by collecting and collating national statistical data that are often subjective, leading to substantial uncertainties in cropland-area estimates, as well as their locations; and (3) extracted from one or more classes of a land use–land cover product in which cropland classes are not the focus of mapping, leading to their mixing with other classes and creating significant errors of omission and commission. These limitations can be overcome by producing high-resolution cropland-extent maps using satellite-sensor data, such as Landsat 30-m resolution or higher. The most fundamental cropland product is the high-resolution cropland-extent map because all higher level cropland products, such as crop-watering method (that is, whether crops are irrigated or rainfed), crop types, cropping intensities, cropland fallows, crop productivity, and crop-water productivity, are dependent on a precise and accurate cropland-extent product. Given these realities, the overarching goal of this study was to produce a Landsat satellite-derived global cropland-extent product at 30-m resolution. The work, which involved a paradigm shift in how global cropland-extent maps are produced, involved the following five key steps: (1) petabyte-scale computing that involved multiyear, 8- to 16-day, time-series Landsat 30-m resolution data for the global land surface; (2) composition of analysis-ready data (ARD) cubes; (3) creation of a large global-reference data hub for machine learning; (4) use of multiple machine-learning algorithms (MLAs) by writing software and computing in the cloud; and (5) Google Earth Engine (GEE) cloud computing. The five key steps involved nine distinct phases. First, the world was segmented into 74 agroecological zones (AEZs). Second, Landsat 8- to 16-day data were used to time-composite 10-band (blue, green, red, near-infrared, short-wave infrared band 1, short-wave infrared band 2, thermal infrared, enhanced vegetation index, normalized difference water index, and normalized difference vegetation index) Landsat 30-m resolution data cubes for every 2- to 4-month time period during 3- to 4-year periods (stated as nominal-year 2015 or, simply, 2015), along with two additional 30-m resolution bands (Shuttle Radar Topography Mission elevation, and slope) in each of the 74 AEZs. Third, more than 100,000 reference-training data samples were collected using ground data (some of which were collected using a mobile application), as well as submeter- to 5-m-resolution, very high-resolution imagery sourced from other reliable sources. Fourth, reference-training data were used to create a knowledge base for separating cropland from noncropland. Fifth, MLAs such as the pixel-based supervised random forest and support-vector machines were written on the GEE using Python and JavaScript. Sixth, object-based recursive hierarchical segmentation algorithm was used, in addition to MLAs, to overcome uncertainties. Seventh, MLAs used the knowledge base to classify and separate cropland from noncropland. Eighth, accuracy assessment was conducted by generating error matrices for each of the 74 AEZs using 19,171 independent validation-data samples. Ninth, cropland areas were computed for all countries of the world and compared with United Nation’s (UN’s) Food and Agricultural Organization (FAO) and other national statistics. The outcome was a Landsat-derived global cropland-extent product at 30-m resolution (GCEP30), which has an overall accuracy of 91.7 percent. For the cropland class, producer’s accuracy was 83.4 percent, and user’s accuracy was 78.3 percent. GCEP30 calculated (using direct pixel count) the global net-cropland area (GNCA) for the year 2015 as 1.873 billion hectares (~12.6 percent of the Earth’s terrestrial area). The continental cropland distribution as a percentage of GNCA was Asia, 33 percent; Europe, 25.5 percent; Africa, 16.7 percent; North America, 14.4 percent; South America, 8.1 percent; and Australia and Oceania, 2.4 percent. The worldwide cropland areas in GCEP30 for 2015 were higher by 236 to 299 million hectares (Mha) compared to national statistics reported elsewhere for the same year (for example, in Food and Agriculture Organization’s corporate statistical database [FAOSTAT] and in the monthly irrigated and rainfed crop areas [MIRCA] database). The global cropland area reported for 2015 increased by 344 Mha (22.5 percent), compared to the year 2000. During the same period (2000–2015), the world’s population increased by 20 percent. Whereas some of these areal increases are real increases in cropland areas, others are due to the types of data, methods, and approaches used. Using the highest known resolution (compared to previous coarse-resolution global products) enabled this study to capture fragmented croplands. Coarse-resolution data compute areas on the basis of subpixels, which, for a large proportion of certain land use–land cover classes, will show only a certain percentage of the total pixel area as actual area. Subpixel areas can lead to substantial uncertainties in area computation, as determining the exact fraction of cropland areas within a coarse-resolution pixel is resource intensive and subject to errors. Other innovations in GCEP30 include reference-data hubs, machine learning, and cloud computing. Cropland areas in 214 countries, territories, departments, and regions were calculated for the year 2015 using GCEP30, on the basis of UN’s global administrative unit layers (GAUL) boundaries. The 10 leading countries in terms of cropland area (as a percentage of the GNCA) were India (9.6 percent), United States (8.95 percent), China (8.82 percent), Russia (8.32 percent), Brazil (3.42 percent), Ukraine (2.32 percent), Canada (2.29 percent), Argentina (2.05 percent), Indonesia (2 percent), and Nigeria (1.91 percent). Together, these 10 countries occupy 50 percent of the global cropland, and they have 52 percent of the global population. Their combined cropland area increased by 2 percent between 2000 and 2015, compared to the substantial increase in population of 517 million (15.5 percent). Together, India, United States, China, and Russia encompass 36 percent of the total area. In the United States and Canada, from 2000 to 2015, cropland decreased by about 2 percent, whereas their populations increased by 14 and 13 percent, respectively. The additional food requirements in these 10 countries, which are caused by increased populations, as well as increasing nutritional demands, are met by production increases in existing cropland or through virtual food trade, or both. More than 18 countries, territories, departments, or regions had 60 percent or more of their geographic area as cropland: Republic of Moldova, San Marino, and Hungary had more than 80 percent of the country’s area as cropland; Denmark, Ukraine, Ireland, and Bangladesh, 70 to 80 percent; and Uruguay, Netherlands, United Kingdom, Spain, Lithuania, Poland, Gaza Strip, Czechia, Italy, India, and Azerbaijan, 60 to 70 percent. Europe and South Asia can be considered agricultural capitals of the world, on the basis of their percentages of geographic area as cropland. United States, China, and Russia, which all have high cropland areas, are ranked second, third, and fourth in the world; India is ranked first. However, the amount of cropland as a percentage of the country’s geographic area is relatively very low for United States (18.3 percent), China (17.7 percent), and Russia (9.5 percent), whereas it is 60.5 percent for India. Most African and South American countries, territories, departments, or regions have less than 15 percent of their geographic area as cropland. China and India together house 36 percent of the world’s population; however, between 2000 and 2015, the amount of China’s cropland area fell by 18.9 percent, owing to urban expansion and the abandonment of farmlands caused by demographic changes (that is, the movement of population from villages to cities). In contrast, China’s population grew by 10 percent. The amount of India’s cropland increased by 8.5 percent, whereas its population grew by 20 percent. This study showed that, out of the 10 leading cropland countries, Ukraine, Nigeria, Russia, and Indonesia showed an 18 to 31 percent increase in cropland areas, on the basis of GCEP30 by the year 2015, compared to 2000. Nigeria’s cropland area increased by 25 percent, and its population increased by 31 percent in the same period. In these countries, food security is maintained by cropland expansion, productivity increases, and virtual food trade. Nevertheless, this trend of increasing net-cropland area and productivity will likely become difficult to maintain, owing to diminishing arable lands and plateauing of 50 years of continual yield increases, requiring policymakers to explore novel and data-supported approaches to solving future food security issues. The GCEP30 product, which can be browsed at full resolution at www.croplands.org , has been released for public download and use through U.S. Geological Survey (USGS)–National Aeronautics and Space Administration (NASA) Land Processes Distributed Active Archive Center (see https://lpdaac.usgs.gov/news/release-of-gfsad-30-meter-cropland-extent-products/ ).

Professional Paper

Using inferential sensors for quality control of Everglades Depth Estimation Network water-level data

The Everglades Depth Estimation Network (EDEN), with over 240 real-time gaging stations, provides hydrologic data for freshwater and tidal areas of the Everglades. These data are used to generate daily water-level and water-depth maps of the Everglades that are used to assess biotic responses to hydrologic change resulting from the U.S. Army Corps of Engineers Comprehensive Everglades Restoration Plan. The generation of EDEN daily water-level and water-depth maps is dependent on high quality real-time data from water-level stations. Real-time data are automatically checked for outliers by assigning minimum and maximum thresholds for each station. Small errors in the real-time data, such as gradual drift of malfunctioning pressure transducers, are more difficult to immediately identify with visual inspection of time-series plots and may only be identified during on-site inspections of the stations. Correcting these small errors in the data often is time consuming and water-level data may not be finalized for several months. To provide daily water-level and water-depth maps on a near real-time basis, EDEN needed an automated process to identify errors in water-level data and to provide estimates for missing or erroneous water-level data. The Automated Data Assurance and Management (ADAM) software uses inferential sensor technology often used in industrial applications. Rather than installing a redundant sensor to measure a process, such as an additional water-level station, inferential sensors, or virtual sensors, were developed for each station that make accurate estimates of the process measured by the hard sensor (water-level gaging station). The inferential sensors in the ADAM software are empirical models that use inputs from one or more proximal stations. The advantage of ADAM is that it provides a redundant signal to the sensor in the field without the environmental threats associated with field conditions at stations (flood or hurricane, for example). In the event that a station does malfunction, ADAM provides an accurate estimate for the period of missing data. The ADAM software also is used in the quality assurance and quality control of the data. The virtual signals are compared to the real-time data, and if the difference between the two signals exceeds a certain tolerance, corrective action to the data and (or) the gaging station can be taken. The ADAM software is automated so that, each morning, the real-time EDEN data are compared to the inferential sensor signals and digital reports highlighting potential erroneous real-time data are generated for appropriate support personnel. The development and application of inferential sensors is easily transferable to other real-time hydrologic monitoring networks.

Florida

The search for reliable aqueous solubility (Sw) and octanol-water partition coefficient (Kow) data for hydrophobic organic compounds; DDT and DDE as a case study

The accurate determination of an organic contaminant’s physico-chemical properties is essential for predicting its environmental impact and fate. Approximately 700 publications (1944–2001) were reviewed and all known aqueous solubilities (S w ) and octanol-water partition coefficients (K ow ) for the organochlorine pesticide, DDT, and its persistent metabolite, DDE were compiled and examined. Two problems are evident with the available database: 1) egregious errors in reporting data and references, and 2) poor data quality and/or inadequate documentation of procedures. The published literature (particularly the collative literature such as compilation articles and handbooks) is characterized by a preponderance of unnecessary data duplication. Numerous data and citation errors are also present in the literature. The percentage of original S w and K ow data in compilations has decreased with time, and in the most recent publications (1994–97) it composes only 6–26 percent of the reported data. The variability of original DDT/DDE S w and K ow data spans 2–4 orders of magnitude, and there is little indication that the uncertainty in these properties has declined over the last 5 decades. A criteria-based evaluation of DDT/DDE S w and K ow data sources shows that 95–100 percent of the database literature is of poor or unevaluatable quality. The accuracy and reliability of the vast majority of the data are unknown due to inadequate documentation of the methods of determination used by the authors. [For example, estimates of precision have been reported for only 20 percent of experimental S w data and 10 percent of experimental K ow data.] Computational methods for estimating these parameters have been increasingly substituted for direct or indirect experimental determination despite the fact that the data used for model development and validation may be of unknown reliability. Because of the prevalence of errors, the lack of methodological documentation, and unsatisfactory data quality, the reliability of the DDT/ DDE S w and K ow database is questionable. The nature and extent of the errors documented in this study are probably indicative of a more general problem in the literature of hydrophobic organic compounds. Under these circumstances, estimation of critical environmental parameters on the basis of S w and K ow (for example, bioconcentration factors, equilibrium partition coefficients) is inadvisable because it will likely lead to incorrect environmental risk assessments. The current state of the database indicates that much greater efforts are needed to: 1) halt the proliferation of erroneous data and references, 2) initiate a coordinated program to develop improved methods of property determination, 3) establish and maintain consistent reporting requirements for physico-chemical property data, and 4) create a mechanism for archiving reliable data for widespread use in the scientific/regulatory community.

Water-Resources Investigations Report

Analysis of ground-water-quality data of the Upper Colorado River basin, water years 1972-92

As part of the U.S. Geological Survey's National Water-Quality Assessment program, an analysis of the existing ground-water-quality data in the Upper Colorado River Basin study unit is necessary to provide information on the historic water-quality conditions. Analysis of the historical data provides information on the availability or lack of data and water-quality issues. The information gathered from the historical data will be used in the design of ground-water-quality studies in the basin. This report includes an analysis of the ground-water data (well and spring data) available for the Upper Colorado River Basin study unit from water years 1972 to 1992 for major cations and anions, metals and selected trace elements, and nutrients. The data used in the analysis of the ground-water quality in the Upper Colorado River Basin study unit were predominantly from the U.S. Geological Survey National Water Information System and the Colorado Department of Public Health and Environment data bases. A total of 212 sites representing alluvial aquifers and 187 sites representing bedrock aquifers were used in the analysis. The available data were not ideal for conducting a comprehensive basinwide water-quality assessment because of lack of sufficient geographical coverage. Evaluation of the ground-water data in the Upper Colorado River Basin study unit was based on the regional environmental setting, which describes the natural and human factors that can affect the water quality. In this report, the ground-water-quality information is evaluated on the basis of aquifers or potential aquifers (alluvial, Green River Formation, Mesaverde Group, Mancos Shale, Dakota Sandstone, Morrison Formation, Entrada Sandstone, Leadville Limestone, and Precambrian) and land-use classifications for alluvial aquifers. Most of the ground-water-quality data in the study unit were for major cations and anions and dissolved-solids concentrations. The aquifer with the highest median concentrations of major ions was the Mancos Shale. The U.S. Environmental Protection Agency secondary maximum contaminant level of 500 milligrams per liter for dissolved solids in drinking water was exceeded in about 75 percent of the samples from the Mancos Shale aquifer. The guideline by the Food and Agriculture Organization of the United States for irrigation water of 2,000 milligrams per liter was also exceeded by the median concentration from the Mancos Shale aquifer. For sulfate, the U.S. Environmental Protection Agency proposed maximum contaminant level of 500 milligrams per liter for drinking water was exceeded by the median concentration for the Mancos Shale aquifer. A total of 66 percent of the sites in the Mancos Shale aquifer exceeded the proposed maximum contaminant level. Metal and selected trace-element data were available for some sites, but most of these data also were below the detection limit. The median concentrations for iron for the selected aquifers and land-use classifications were below the U.S. Environmental Protection Agency secondary maximum contaminant level of 300 micrograms per liter in drinking water. Median concentration of manganese for the Mancos Shale exceeded the U.S. Environmental Protection Agency secondary maximum contaminant level of 50 micrograms per liter in drinking water. The highest selenium concentrations were in the alluvial aquifer and were associated with rangeland. However, about 22 percent of the selenium values from the Mancos Shale exceeded the U.S. Environmental Protection Agency maximum contaminant level of 50 micrograms per liter in drinking water. Few nutrient data were available for the study unit. The only nutrient species presented in this report were nitrate-plus-nitrite as nitrogen and orthophosphate. Median concentrations for nitrate-plus-nitrite as nitrogen were below the U.S. Environmental Protection Agency maximum contaminant level of 10 milligrams per liter in drinking water except for 0.02 percent of the sites in the alluvial aquifer and 0.03 percent of the sites in the Mancos Shale. Concentrations of orthophosphate did not vary significantly among aquifers or land-use classifications. Historic water-quality data from wells and springs helped to characterize the regional distribution of ground-water quality information in the Upper Colorado River Basin study unit. The historical ground-water data summarized in this report will be used in the design of a ground-water-quality network. Because ground-water-quality issues in the study unit are related to high dissolved solids, sulfate, selenium, and nutrients, this report discusses some of the important findings related to these issues.

Colorado

On the documentation, independence, and stability of widely used seismological data products

Earthquake scientists have traditionally relied on relatively small data sets recorded on small numbers of instruments. With advances in both instrumentation and computational resources, the big-data era, including an established norm of open data-sharing, allows seismologists to explore important issues using data volumes that would have been unimaginable in earlier decades. Alongside with these developments, the community has moved towards routine production of interpreted data products such as seismic moment tensor catalogs that have provided an additional boon to earthquake science. As these products have become increasingly familiar and useful, it is important to bear in mind that they are not data, but rather interpreted data products. As such, they differ from data in ways that can be important, but not always appreciated. Important - and sometimes surprising - issues can arise if methodology is not fully described, data from multiple sources are included, or data products are not versioned (time-stamped). The line between data and data products is sometimes blurred, leading to an underappreciation of issues that affect data products. This note illustrates examples from two widely used data products: moment tensor catalogs and Did You Feel It? (DYFI) macroseismic intensity values. These examples show that increasing a data product’s documentation, independence, and stability can make it even more useful. To ensure the reproducibility of studies using data products, time-stamped products should be preserved, for example as electronic supplements to published papers, or, ideally, a more permanent repository.

Frontiers in Earth Science

Monitoring landscape dynamics in central U.S. grasslands with harmonized Landsat-8 and Sentinel-2 time series data

Remotely monitoring changes in central U.S. grasslands is challenging because these landscapes tend to respond quickly to disturbances and changes in weather. Such dynamic responses influence nutrient cycling, greenhouse gas contributions, habitat availability for wildlife, and other ecosystem processes and services. Traditionally, coarse-resolution satellite data acquired at daily intervals have been used for monitoring. Recently, the harmonized Landsat-8 and Sentinel-2 (HLS) data increased the temporal frequency of the data. Here we investigated if the increased data frequency provided adequate observations to characterize highly dynamic grassland processes. We evaluated HLS data available for 2016 to (1) determine if data from Sentinel-2 contributed to an improvement in characterizing landscape processes over Landsat-8 data alone, and (2) quantify how observation frequency impacted results. Specifically, we investigated into estimating annual vegetation phenology, detecting burn scars from fire, and modeling within-season wetland hydroperiod and growth of aquatic vegetation. We observed increased sensitivity to the start of the growing season (SOST) with the HLS data. Our estimates of the grassland SOST compared well with ground estimates collected at a phenological camera site. We used the Continuous Change Detection and Classification (CCDC) algorithm to assess if the HLS data improved our detection of burn scars following grassland fires and found that detection was considerably influenced by the seasonal timing of the fires. The grassland burned in early spring recovered too quickly to be detected as change events by CCDC; instead, the spectral characteristics following these fires were incorporated as part of the ongoing time-series models. In contrast, the spectral effects from late-season fires were detected both by Landsat-8 data and HLS data. For wetland-rich areas, we used a modified version of the CCDC algorithm to track within-season dynamics of water and aquatic vegetation. The addition of Sentinel-2 data provided the potential to build full time series models to better distinguish different wetland types, suggesting that the temporal density of data was sufficient for within-season characterization of wetland dynamics. Although the different data frequency, in both the spatial and temporal dimensions, could cause inconsistent model estimation or sensitivity sometimes; overall, the temporal frequency of the HLS data improved our ability to track within-season grassland dynamics and improved results for areas prone to cloud contamination. The results suggest a greater frequency of observations, such as from harmonizing data across all comparable Landsat and Sentinel sensors, is still needed. For our study areas, at least a 3-day revisit interval during the early growing season (weeks 14–17) is required to provide a >50% probability of obtaining weekly clear observations.

Minnesota