Geology ReportsSearch

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Data management system for USGS/USEPA urban hydrology studies program

A data management system was developed to store, update, and retrieve data collected in urban stormwater studies jointly conducted by the U.S. Geological Survey and U.S. Environmental Protection Agency in 11 cities in the United States. The data management system is used to retrieve and combine data from USGS data files for use in rainfall, runoff, and water-quality models and for data computations such as storm loads. The system is based on the data management aspect of the Statistical Analysis System (SAS) and was used to create all the data files in the data base. SAS is used for storage and retrieval of basin physiography, land-use, and environmental practices inventory data. Also, storm-event water-quality characteristics are stored in the data base. The advantages of using SAS to create and manage a data base are many with a few being that it is simple, easy to use, contains a comprehensive statistical package, and can be used to modify files very easily. Data base system development has progressed rapidly during the last two decades and the data managment system concepts used in this study reflect the advancement made in computer technology during this era. Urban stormwater data is, however, just one application for which the system can be used. (USGS)

Open-File Report

Summary of selected U.S. Geological survey data on domestic well water quality for the Centers for Disease Control's National Environmental Public Health Tracking Program

About 10 to 30 percent of the population in most States uses domestic (private) water supply. In many States, the total number of people served by domestic supplies can be in the millions. The water quality of domestic supplies is inconsistently regulated and generally not well characterized. The U.S. Geological Survey (USGS) has two water-quality data sets in the National Water Information System (NWIS) database that can be used to help define the water quality of domestic-water supplies: (1) data from the National Water-Quality Assessment (NAWQA) Program, and (2) USGS State data. Data from domestic wells from the NAWQA Program were collected to meet one of the Program's objectives, which was to define the water quality of major aquifers in the United States. These domestic wells were located primarily in rural areas. Water-quality conditions in these major aquifers as defined by the NAWQA data can be compared because of the consistency of the NAWQA sampling design, sampling protocols, and water-quality analyses. The NWIS database is a repository of USGS water data collected for a variety of projects; consequently, project objectives and analytical methods vary. This variability can bias statistical summaries of contaminant occurrence and concentrations; nevertheless, these data can be used to define the geographic distribution of contaminants. Maps created using NAWQA and USGS State data in NWIS can show geographic areas where contaminant concentrations may be of potential human-health concern by showing concentrations relative to human-health water-quality benchmarks. On the basis of national summaries of detection frequencies and concentrations relative to U.S. Environmental Protection Agency (USEPA) human-health benchmarks for trace elements, pesticides, and volatile organic compounds, 28 water-quality constituents were identified as contaminants of potential human-health concern. From this list, 11 contaminants were selected for summarization of water-quality data in 16 States (grantee States) that were funded by the Environmental Public Health Tracking (EPHT) Program of the Centers for Disease Control and Prevention (CDC). Only data from domestic-water supplies were used in this summary because samples from these wells are most relevant to human exposure for the targeted population. Using NAWQA data, the concentrations of the 11 contaminants were compared to USEPA human-health benchmarks. Using NAWQA and USGS State data in NWIS, the geographic distribution of the contaminants were mapped for the 16 grantee States. Radon, arsenic, manganese, nitrate, strontium, and uranium had the largest percentages of samples with concentrations greater than their human-health benchmarks. In contrast, organic compounds (pesticides and volatile organic compounds) had the lowest percentages of samples with concentrations greater than human-health benchmarks. Results of data retrievals and spatial analysis were compiled for each of the 16 States and are presented in State summaries for each State. Example summary tables, graphs, and maps based on USGS data for New Jersey are presented to illustrate how USGS water-quality and associated ancillary geospatial data can be used by the CDC to address goals and objectives of the EPHT Program.

Scientific Investigations Report

Determination of study reporting limits for pesticide constituent data for the California Groundwater Ambient Monitoring and Assessment Program Priority Basin Project, 2004–2018—Part 1: National Water Quality Schedules 2003, 2032, or 2033, and 2060

The California Groundwater Ambient Monitoring and Assessment Program Priority Basin Project (GAMA-PBP) is a long-term cooperative project designed to assess the quality of groundwater resources used for public and domestic drinking water supplies in the State of California, to monitor and evaluate changes to that quality, to investigate the human and natural factors controlling water quality, and to improve the availability of comprehensive groundwater quality data and information. Between May 18, 2004, and May 3, 2018, the GAMA-PBP collected 3001 groundwater samples for analysis of pesticide constituents by the U.S. Geological Survey (USGS) National Water Quality Laboratory (NWQL)(note that ‘pesticide constituents’ includes parent compounds and degradates). Of these samples, 2994 were analyzed for pesticide constituents on schedules 2003, 2032, or 2033 (65 to 84 constituents), and 840 were analyzed for pesticide constituents on schedule 2060 (58 constituents). The original dataset reported by the NWQL to the USGS National Water Information System (NWIS) database contained a total of 2,688 detections of 78 pesticide constituents and 253,825 non-detections. In this original dataset, 33 percent of the 3,001 samples analyzed had reported detections of one or more pesticide constituents. This report describes the GAMA-PBP data-quality objectives for pesticide data, the procedures used to establish study reporting limits, and use of those reporting limits to censor the data from the NWQL so that the final data published by the GAMA-PBP meet these data-quality objectives. The final GAMA-PBP dataset for samples collected from May 2004 to May 2018, after censoring, had a total of 1,632 detections of 37 pesticide constituents. In the final GAMA-PBP dataset, 25 percent of the 3,001 samples analyzed had detections of one or more pesticide constituents. The presence of pesticides in groundwater is commonly evaluated by calculating detection frequencies. Detection frequencies for pesticides are sensitive to detection limits and method performance for concentrations near those limits; therefore, the two primary data quality issues addressed in the GAMA-PBP data-quality objectives for pesticides are (1) establishing criteria for classifying data from the laboratory as detections or non-detections for the purpose of data reporting by the project and (2) accounting for changes in analytical methods or method performance over time. The GAMA-PBP addresses these issues by developing study reporting limits that are used as the boundary between detections and non-detections for the reporting of GAMA-PBP results. These reporting limits are defined from method detection limits (MDLs) provided by the NWQL, unless examination of results from laboratory set blanks (LSBs) and GAMA-PBP field blanks indicates that a higher concentration censoring limit is warranted. The GAMA-PBP selected the MDL as the primary choice for defining study reporting limits for consistency with U.S. Environmental Protection Agency (EPA) guidelines for reporting detections of pesticides and other organic constituents. A five-step procedure is used to develop study reporting limits and censor the GAMA-PBP dataset accordingly. The effect of the censoring at each step is described to provide information about the relative effect of each step on the overall censoring of the dataset. Steps 1 and 2 can be implemented at the time the data are received, whereas steps 3−5 require information accumulated over an extended period. Step 1: Reject results that were most likely the result of specific contamination instances attributable to unusual field or laboratory conditions during sample collection or processing. Two such instances were identified, leading to rejection of 25 detections, which were assigned a data-quality indicator code of “Q” for “reviewed and rejected” in the NWIS database. Step 2: Use the NWQL MDLs in effect at the time each sample was analyzed as the reporting limit. A total of 506 detections were censored on this basis. Step 3: Use the maximum MDL established by the NWQL during July 2004–August 2018 (MDLmax) as the reporting limit. The rationale for using the MDLmax as the reporting limit is based primarily on the observation that the concentrations of MDLs generally increased over time. A total of 438 detections were censored on this basis. Step 4: Use the LSBs to identify periods of greater potential laboratory contamination bias and define raised reporting limits to be used during those periods. These periods were defined by using a moving average detection frequency approach. For consistency with the NWQL procedures for defining raised reporting limits on the basis of detections in LSBs, the raised reporting limits were defined as equal to three times the highest concentration measured in an LSB during the period. A total of 25 detections in groundwater samples analyzed during periods of increased laboratory contamination bias were censored. Step 5: Use the LSBs and field blanks to identify potential contamination bias from field or laboratory processes outside of the time periods identified in step 4. The NWQL protocols were used to define the MDLs from blanks analyzed outside of the periods identified in step 4. If an MDL defined from blanks was greater than the MDLmax, the MDL defined from blanks was used to censor the data. One constituent had a study reporting limit defined on this basis, and a total of 62 detections in groundwater samples were censored. As of 2019, the USGS NWIS database does not have the capability to store both the original value reported by the NWQL and the final value published by the GAMA-PBP that reflects application of the quality-control censoring described in this report. In the interim, while this capability is developed, the 1,031 results censored in steps 2−5 are blocked from public release in NWIS, and the GAMA-PBP has published the original and final values in a USGS data release accompanying this report. The entire GAMA-PBP final dataset for pesticide constituents on schedules 2003, 2032, or 2033, or on schedule 2060 is publicly available in that USGS data release, through the USGS GAMA-PBP public web portal, and through the California State Water Resources Control Board GAMA public groundwater information system.

California

Evaluating hydrologic data products for scientific and management applications related to potential future streamflow conditions in the Upper Mississippi and Illinois Rivers

The hydrology of the Upper Mississippi and Illinois Rivers is a fundamental driver of ecosystem patterns and processes across a large portion of the United States. Quantitative hydrologic data for the main stems of these rivers underlie numerous scientific investigations, statistical models, and decision-making processes for local, State, and Federal agencies involved in the Upper Mississippi River Restoration program. Although historical hydrologic data exist, data representing potential future conditions of the Upper Mississippi and Illinois Rivers lack the resolution necessary to anticipate biotic and abiotic responses to altered hydrology and to determine resilient management actions. A source of future hydrologic scenarios is the readily available LOCA–VIC–mizuRoute hydrologic data products (named for the chain of models the data are produced from—localized constructed analogs, Variable Infiltration Capacity macroscale hydrological model, and the mizuRoute hydrologic routing model—that we shorten further to LVM in this report) that include simulated discharges for historic and future timeframes. The objective of this study is to assess the reliability of the hydrologic data products for their use in Upper Mississippi River Restoration program applications. Key study questions are (1) do the hydrologic data products reproduce characteristics of hydrology necessary to support ecological modeling and restoration decision-making applications within the Upper Mississippi River Restoration program? and (2) are there geographic differences in the reliability of the hydrologic data products? Seven characteristics of river hydrology were selected related to flow magnitude, seasonality, and regime for evaluation. The seven characteristics were calculated using observed and historical simulated hydrologic data at 19 U.S. Geological Survey streamgages throughout the basins of the Upper Mississippi and Illinois Rivers; two streamgages are located on the main stem of the Mississippi River and two streamgages are located on the main stem of the Illinois River. Statistical comparisons between observed and historical simulated characteristics indicated that the hydrologic data products did not reliably represent historical hydrologic conditions in the basin or main stem. The hydrologic data products we evaluated could not reliably capture the overall hydrologic regime or flow magnitudes; the latter is evidenced by substantial underestimates of discharge at most streamgages. Seasonal hydrologic characteristics were captured more reliably than flow magnitude, but overall correspondence was low for most streamgages. A weak latitudinal pattern in seasonal characteristics indicated the hydrologic data products poorly represent streamflow timing in snow-affected regions of the basin. Discrepancies in magnitude, seasonality, and regime indicate the potential for multiple sources of error. Because poor correspondence was present across all 19 streamgages, it was not possible to identify specific drivers of poor performance (that is, drainage area or geography). The modeling chain should be evaluated for biases associated with meteorologic forcing data, as well as hydrologic model formulation and calibration. We conclude that the hydrologic data products we evaluated appear unsuitable for applications tied to habitat and ecosystem restoration and management in the Upper Mississippi and Illinois Rivers. Plans to develop a future hydrology dataset for the Upper Mississippi River Restoration program would benefit from ongoing work to improve global climate model output downscaling methods, to improve hydrologic models, to make use of innovations in machine-learning approaches for projecting hydrology, and other efforts. The framework developed herein to evaluate hydrometeorological outputs generated using global climate models for a specific water resources application is a transferrable approach that could be applied to other data products and river systems.

Illinois, Indiana, Iowa, Minnesota, Missouri, Sout

Effect of mineral deposit data on predictions from the three-part approach to quantitative mineral resource assessment—A study of 16 previous U.S. Geological Survey assessments

The three-part approach to quantitative mineral resource assessment requires information about the properties of undiscovered mineral deposits in an assessment area. These properties are unknown, so the properties of discovered mineral deposits of the same mineral deposit type are used instead. In the three-part approach, these discovered mineral deposits come from around the world, and their properties constitute the pooled data for that mineral deposit type. Alternatively, these discovered mineral deposits could come from the assessment area, and their properties constitute the tract data for that mineral deposit type. Tract data may be more representative of the undiscovered mineral deposits in the assessment area than the pooled data. The goal of this study was to determine whether resource predictions using pooled data are equivalent to resource predictions using tract data. To this end, 16 previous U.S Geological Survey assessments were studied. For each assessment, resources were predicted for one undiscovered mineral deposit in the assessment area. One set of predictions used pooled data, and another used tract data. The two sets of predictions were compared with an equivalence test, using the six assessment statistics that are commonly reported for mineral resource assessments. Practical equivalence is the condition that two corresponding assessment statistics are within a factor of 1.5 of one another. For each of 2 assessments, all 6 assessment statistics were practically equivalent. For both assessments, the assessment statistics from the pooled data, relative to the corresponding assessment statistics from the tract data, ranged from 1.30 times smaller to 1.03 times larger. For each of 14 assessments, 1 or more of the 6 assessment statistics were not practically equivalent. The assessment statistics from the pooled data, relative to the corresponding assessment statistics from the tract data, ranged from 26.6 times smaller to 5.53 times larger. The use of pooled data has been a standard procedure in the three-part approach since at least 1986. The 16 assessments in this study are not a representative sample of those prior assessments that used pooled data. So, it is inappropriate to use the study results to infer whether pooled data affected the resource predictions for those prior assessments.

Scientific Investigations Report

Geometric quality assessment of lidar data based on swath overlap

This paper provides guidelines on quantifying the relative horizontal and vertical errors observed between conjugate features in the overlapping regions of lidar data. The quantification of these errors is important because their presence quantifies the geometric quality of the data. A data set can be said to have good geometric quality if measurements of identical features, regardless of their position or orientation, yield identical results. Good geometric quality indicates that the data are produced using sensor models that are working as they are mathematically designed, and data acquisition processes are not introducing any unforeseen distortion in the data. High geometric quality also leads to high geolocation accuracy of the data when the data acquisition process includes coupling the sensor with geopositioning systems. Current specifications (e.g. Heidemann 2014) do not provide adequate means to quantitatively measure these errors, even though they are required to be reported. Current accuracy measurement and reporting practices followed in the industry and as recommended by data specification documents also potentially underestimate the inter-swath errors, including the presence of systematic errors in lidar data. Hence they pose a risk to the user in terms of data acceptance (i.e. a higher potential for Type II error indicating risk of accepting potentially unsuitable data). For example, if the overlap area is too small or if the sampled locations are close to the center of overlap, or if the errors are sampled in flat regions when there are residual pitch errors in the data, the resultant Root Mean Square Differences (RMSD) can still be small. To avoid this, the following are suggested to be used as criteria for defining the inter-swath quality of data: a) Median Discrepancy Angle b) Mean and RMSD of Horizontal Errors using DQM measured on sloping surfaces c) RMSD for sampled locations from flat areas (defined as areas with less than 5 degrees of slope) It is suggested that 4000-5000 points are uniformly sampled in the overlapping regions of the point cloud, and depending on the surface roughness, to measure the discrepancy between swaths. Care must be taken to sample only areas of single return points only. Point-to-Plane distance based data quality measures are determined for each sample point. These measurements are used to determine the above mentioned parameters. This paper details the measurements and analysis of measurements required to determine these metrics, i.e. Discrepancy Angle, Mean and RMSD of errors in flat regions and horizontal errors obtained using measurements extracted from sloping regions (slope greater than 10 degrees). The research is a result of an ad-hoc joint working group of the US Geological Survey and the American Society for Photogrammetry and Remote Sensing (ASPRS) Airborne Lidar Committee.

Conference Paper

Narrowing historical uncertainty: probabilistic classification of ambiguously identified tree species in historical forest survey data

Historical data have increasingly become appreciated for insight into the past conditions of ecosystems. Uses of such data include assessing the extent of ecosystem change; deriving ecological baselines for management, restoration, and modeling; and assessing the importance of past conditions on the composition and function of current systems. One historical data set of this type is the Public Land Survey (PLS) of the United States General Land Office, which contains data on multiple tree species, sizes, and distances recorded at each survey point, located at half-mile (0.8 km) intervals on a 1-mi (1.6 km) grid. This survey method was begun in the 1790s on US federal lands extending westward from Ohio. Thus, the data have the potential of providing a view of much of the US landscape from the mid-1800s, and they have been used extensively for this purpose. However, historical data sources, such as those describing the species composition of forests, can often be limited in the detail recorded and the reliability of the data, since the information was often not originally recorded for ecological purposes. Forest trees are sometimes recorded ambiguously, using generic or obscure common names. For the PLS data of northern Wisconsin, USA, we developed a method to classify ambiguously identified tree species using logistic regression analysis, using data on trees that were clearly identified to species and a set of independent predictor variables to build the models. The models were first created on partial data sets for each species and then tested for fit against the remaining data. Validations were conducted using repeated, random subsets of the data. Model prediction accuracy ranged from 81% to 96% in differentiating congeneric species among oak, pine, ash, maple, birch, and elm. Major predictor variables were tree size, associated species, landscape classes indicative of soil type, and spatial location within the study region. Results help to clarify ambiguities formerly present in maps of historic ecosystems for the region and can be applied to PLS datasets elsewhere, as well as other sources of ambiguous historical data. Mapping the newly classified data with ecological land units provides additional information on the distribution, abundance, and associations of tree species, as well as their relationships to environmental gradients before the industrial period, and clarifies the identities of species formerly mapped only to genus. We offer some caveats on the appropriate use of data derived in this way, as well as describing their potential.

Ecosystems

Archive of bathymetry data collected in South Florida from 1995 to 2015

Description Land development and alterations of the ecosystem in south Florida over the past 100 years have decreased freshwater and increased nutrient flows into many of Florida's estuaries, bays, and coastal regions. As a result, there has been a decrease in the water quality in many of these critical habitats, often prompting seagrass die-offs and reduced fish and aquatic life populations. Restoration of water quality in many of these habitats will depend partly upon using numerical-circulation and sediment-transport models to establish water-quality targets and to assess progress toward reaching restoration targets. Application of these models is often complicated because of complex sea floor topography and tidal flow regimes. Consequently, accurate and modern sea-floor or bathymetry maps are critical for numerical modeling research. Modern bathymetry data sets will also permit a comparison to historical data in order to help assess sea-floor changes within these critical habitats. New and detailed data sets also support marine biology studies to help understand migratory and feeding habitats of marine life. This data series is a compilation of 13 mapping projects conducted in south Florida between 1995 and 2015 and archives more than 45 million bathymetric soundings. Data were collected primarily with a single beam sound navigation and ranging (sonar) system called SANDS developed by the U.S. Geological Survey (USGS) in 1993. Bathymetry data for the Estero Bay project were supplemented with the National Aeronautics and Space Administration's (NASA) Experimental Advanced Airborne Research Lidar (EAARL) system. Data from eight rivers in southwest Florida were collected with an interferometric swath bathymetry system. The projects represented in this data series were funded by the USGS Coastal and Marine Geology Program (CMGP), the USGS South Florida Ecosystem Restoration Project- formally named Placed Based Studies, and other non-Federal agencies. The purpose of the data collection for all these projects was to support one or more of the following scientific aspects: numerical model applications, sea floor change analysis, or marine habitat investigations. This report serves as an archive of processed bathymetry sounding data, digital bathymetric contours, digital bathymetric maps, sea floor surface grids, and formal Federal Geographic Data Committee (FGDC) metadata. Refer to the Abbreviations page for explanations of acronyms and abbreviations used in this report. Since 2006, the USGS St. Petersburg Coastal and Marine Science Center (SPCMSC) assigns a unique identifier or Field Activity Number (FAN) for each field data collection. Projects described in this report conducted prior to 2006 do not have a FAN. Data from the 13 projects presented in this report provided critical hydrographic information to support multiple science projects in south Florida. The projects and the types of sounding data collected are: Florida Bay (1995-1999) - single-beam Lake Okeechobee (2001) - single-beam Tampa Bay (2001-2004) - single-beam Caloosahatchee River (2002)- single-beam Estero Bay to Matlacha Pass and offshore to Wiggins Pass (2003) - single-beam and airborne lidar North and Northwest Forks of the Loxahatchee and Lower St. Lucie Rivers (2003) - single-beam South Charlotte Harbor and offshore Sanibel Island (2003-2004) - single-beam Shark River and Trout Creek (2004) - single-beam and interferometric swath Southwest Florida Rivers (2004) - interferometric swath Offshore from Wiggins Pass to Cape Romano (2005) - single-beam Ten Thousand Islands (2009) - single-beam Lemon Bay (2011) - single-beam Southwest Florida Rivers (2015) - interferometric swath

Florida

Post-disaster supply chain interdependent critical infrastructure system restoration: A review of data necessary and available for modeling

The majority of restoration strategies in the wake of large-scale disasters have focused on short-term emergency response solutions. Few consider medium- to long-term restoration strategies to reconnect urban areas to national supply chain interdependent critical infrastructure systems (SCICI). These SCICI promote the effective flow of goods, services, and information vital to the economic vitality of an urban environment. To re-establish the connectivity that has been broken during a disaster between the different SCICI, relationships between these systems must be identified, formulated, and added to a common framework to form a system-level restoration plan. To accomplish this goal, a considerable collection of SCICI data is necessary. The aim of this paper is to review what data are required for model construction, the accessibility of these data, and their integration with each other. While a review of publicly available data reveals a dearth of real-time data to assist modeling long-term recovery following an extreme event, a significant amount of static data does exist and these data can be used to model the complex interdependencies needed. For the sake of illustration, a particular SCICI (transportation) is used to highlight the challenges of determining the interdependencies and creating models capable of describing the complexity of an urban environment with the data publicly available. Integration of such data as is derived from public domain sources is readily achieved in a geospatial environment, after all geospatial infrastructure data are the most abundant data source and while significant quantities of data can be acquired through public sources, a significant effort is still required to gather, develop, and integrate these data from multiple sources to build a complete model. Therefore, while continued availability of high quality, public information is essential for modeling efforts in academic as well as government communities, a more streamlined approach to a real-time acquisition and integration of these data is essential.

Data Science Journal

Statistical Approaches to Interpretation of Local, Regional, and National Highway-Runoff and Urban-Stormwater Data

Decision makers need viable methods for the interpretation of local, regional, and national-highway runoff and urban-stormwater data including flows, concentrations and loads of chemical constituents and sediment, potential effects on receiving waters, and the potential effectiveness of various best management practices (BMPs). Valid (useful for intended purposes), current, and technically defensible stormwater-runoff models are needed to interpret data collected in field studies, to support existing highway and urban-runoffplanning processes, to meet National Pollutant Discharge Elimination System (NPDES) requirements, and to provide methods for computation of Total Maximum Daily Loads (TMDLs) systematically and economically. Historically, conceptual, simulation, empirical, and statistical models of varying levels of detail, complexity, and uncertainty have been used to meet various data-quality objectives in the decision-making processes necessary for the planning, design, construction, and maintenance of highways and for other land-use applications. Water-quality simulation models attempt a detailed representation of the physical processes and mechanisms at a given site. Empirical and statistical regional water-quality assessment models provide a more general picture of water quality or changes in water quality over a region. All these modeling techniques share one common aspect-their predictive ability is poor without suitable site-specific data for calibration. To properly apply the correct model, one must understand the classification of variables, the unique characteristics of water-resources data, and the concept of population structure and analysis. Classifying variables being used to analyze data may determine which statistical methods are appropriate for data analysis. An understanding of the characteristics of water-resources data is necessary to evaluate the applicability of different statistical methods, to interpret the results of these techniques, and to use tools and techniques that account for the unique nature of water-resources data sets. Populations of data on stormwater-runoff quantity and quality are often best modeled as logarithmic transformations. Therefore, these factors need to be considered to form valid, current, and technically defensible stormwater-runoff models. Regression analysis is an accepted method for interpretation of water-resources data and for prediction of current or future conditions at sites that fit the input data model. Regression analysis is designed to provide an estimate of the average response of a system as it relates to variation in one or more known variables. To produce valid models, however, regression analysis should include visual analysis of scatterplots, an examination of the regression equation, evaluation of the method design assumptions, and regression diagnostics. A number of statistical techniques are described in the text and in the appendixes to provide information necessary to interpret data by use of appropriate methods. Uncertainty is an important part of any decisionmaking process. In order to deal with uncertainty problems, the analyst needs to know the severity of the statistical uncertainty of the methods used to predict water quality. Statistical models need to be based on information that is meaningful, representative, complete, precise, accurate, and comparable to be deemed valid, up to date, and technically supportable. To assess uncertainty in the analytical tools, the modeling methods, and the underlying data set, all of these components need be documented and communicated in an accessible format within project publications.

Open-File Report

Quality-assurance data for routine water analyses by the U.S. Geological Survey Laboratory in Troy, New York—July 1995 through June 1997

The laboratory for analysis of low-ionic-strength water at the U.S. Geological Survey (USGS) office in Troy, N.Y. analyzes samples collected by USGS projects in the Northeast. The laboratory’s quality-assurance program is based on internal and interlaboratory quality-assurance samples and quality-control procedures developed to ensure proper sample collection, processing, and analysis. For the time period addressed in this report, the quality-assurance/quality-control data were stored in the laboratory’s SAS data-management system, which provides efficient review, compilation, and plotting of quality-assurance/quality-control data. This report presents and discusses samples analyzed from July 1995 through June 1997. Quality-control results for 19 analytical procedures were evaluated for bias and precision. Control charts show that data from ten of the analytical procedures were biased throughout the analysis period for either high-concentration or low-concentration samples but were within control limits; these procedures were: acid-neutralizing capacity, total monomeric aluminum, ammonium, calcium, chloride, dissolved organic carbon, magnesium, nitrate (ion chromatography), nitrate (colorimetric method), and sulfate. Four of the analytical procedures were occasionally biased but were within control limits; they were: fluoride, pH, silicon, and sodium. Results from the filter-blank and analytical-blank analyses indicate that all analytical procedures in which blanks were run were within control limits, although values for a few blanks were outside the control limits. Sampling and analysis precision are evaluated herein in terms of the coefficient of variation obtained for triplicate samples in 14 of the 19 procedures. Data-quality objectives (DQO’s) were met by at least 92 percent of the samples analyzed in all procedures except acid neutralizing capacity (80 percent of samples met objectives), total monomeric aluminum (87 percent of samples met objectives), organic monomeric aluminum (89 percent of samples met objectives), and chloride (89 percent of samples met objectives). The data are insufficient to evaluate the DQO’s for total aluminum. Results of the USGS interlaboratory Standard Reference Sample Program indicated acceptable data quality for most constituents over the time period. The results of the P-sample (low-ionic strength constituent) analysis indicated high data quality with good ratings in all studies. The T-sample (trace constituent) had unacceptable ratings in two studies, but received satisfactory ratings in the others. The N-sample (nutrient constituent) studies had an unacceptable rating in one and an excellent rating in the other. Environment Canada’s NWRI program results indicated that at least 90 percent of the samples met data-quality objectives in 9 of the 12 analyses; exceptions were calcium, chloride, and silicon. Data-quality objectives were not met for calcium samples in two NWRI studies, but all of the samples analyzed were within control limits for the remaining studies. Data-quality objectives were not met for 32 percent of samples analyzed for chloride and 27 percent of samples analyzed for silicon. Results from blind reference-sample analyses indicated that data-quality objectives were met by at least 90 percent of the calcium, pH, potassium, and sodium samples. Data-quality objectives were met by 77 percent of the chloride samples, 83 percent of the magnesium samples, and 80 percent of the sulfate samples. There is insufficient data to evaluate the specific conductance samples.

Open-File Report

Water-quality data-collection activities in Colorado and Ohio; Phase II, Evaluation of 1984 field and laboratory quality-assurance practices

Serious questions have been raised by Congress about the usefulness of water-quality data for addressing issues of regional and national scope and, especially, for characterizing the current quality of the Nation's streams and ground water. In response, the U.S. Geological Survey has undertaken a pilot study in Colorado and Ohio to (1) determine the characteristics of current (1984) water-quality data-collection activities of Federal, regional, State, and local agencies, and academic institutions; and (2) determine how well the data from these activities, collected for various purposes and using different procedures, can be used to improve our ability to answer major broad-scope questions, such as: A. What are (or were) natural or near-natural water-quality conditions? B. What are existing water-quality conditions? C. How has water quality changed, and how do the changes relate to human activities? Colorado and Ohio were chosen for the pilot study largely because they represent regions with different types of waterquality concerns and programs. The study has been divided into three phases, the objectives of which are: Phase I--Inventory water-quality data-collection programs, including costs, and identify those programs that met a set of broad criteria for producing data that are potentially appropriate for water-quality assessments of regional and national scope. Phase II--Evaluate the quality assurance of field and laboratory procedures used in producing the data from programs that met the broad criteria of Phase I. Phase III--Compile the qualifying data and evaluate the adequacy of this data base for addressing selected water-quality questions of regional and national scope. Water-quality data are collected by a large number of organizations for diverse purposes ranging from meeting statutory requirements to research on water chemistry. Combining these individual data bases is an appealing and potentially cost-effective way to attempt to develop a data base adequate for regional or national water-quality assessments. However, to combine data from diverse sources, field and laboratory procedures used to produce the data need to be equivalent and need to meet specific qualityassurance standards. It is these factors that are the focus of Phase II, which is described in this report. In the first phase of this study, an inventory was made of all public organizations and academic institutions that undertook water-quality data-collection activities in Colorado and Ohio in 1984. Water-quality programs identified in Phase I were tested against a set of broad screening criteria. A total of 44 waterquality programs in Colorado and 29 programs in Ohio passed the Phase-I screen and were examined in Phase II. These programs accounted for an estimated 165,000 analyses in Colorado and 76,300 analyses in Ohio for 20 selected constituents and properties. Although qualifying programs included both surface- and ground-water sampling, they emphasized surface waters and produced few groundwater analyses (3,660 for Colorado and 470 for Ohio). For Phase II, information about field and laboratory qualityassurance practices was provided by each organization and its supporting laboratories through questionnaires. This information was evaluated against a set of specific criteria for field and laboratory practices. The criteria were developed from guidelines published by public agencies and professional organizations such as the American Public Health Association, the U.Sc, Environmental Protection Agency, and the U.S. Geological Survey. Each of the eight criteria that comprise the Phase-II screen fall into one of two major categories--field practices or laboratory practices.

Colorado, Ohio

Definitions of database files and fields of the Personal Computer-Based Water Data Sources Directory

This report describes the data-base files and fields of the personal computer-based Water Data Sources Directory (WDSD). The personal computer-based WDSD was derived from the U.S. Geological Survey (USGS) mainframe computer version. The mainframe version of the WDSD is a hierarchical data-base design. The personal computer-based WDSD is a relational data- base design. This report describes the data-base files and fields of the relational data-base design in dBASE IV (the use of brand names in this abstract is for identification purposes only and does not constitute endorsement by the U.S. Geological Survey) for the personal computer. The WDSD contains information on (1) the type of organization, (2) the major orientation of water-data activities conducted by each organization, (3) the names, addresses, and telephone numbers of offices within each organization from which water data may be obtained, (4) the types of data held by each organization and the geographic locations within which these data have been collected, (5) alternative sources of an organization's data, (6) the designation of liaison personnel in matters related to water-data acquisition and indexing, (7) the volume of water data indexed for the organization, and (8) information about other types of data and services available from the organization that are pertinent to water-resources activities.

Open-File Report

National Geochemical Database: Reformatted data from the National Uranium Resource Evaluation (NURE) Hydrogeochemical and Stream Sediment Reconnaissance (HSSR) program

The National Uranium Resource Evaluation (NURE) Hydrogeochemical and Stream Sediment Reconnaissance (HSSR) program produced a large amount of geochemical data. To fully understand how these data were generated, it is recommended that you read the History of NURE HSSR Program for a summary of the entire program. By the time the NURE program had ended, the HSSR data consisted of 894 separate data files stored with 47 different formats. Many files contained duplication of data found in other files. The University of Oklahoma's Information Systems Programs of the Energy Resources Institute (ISP) was contracted by the Department of Energy to enhance the accessibility and usefulness of the NURE HSSR data. ISP created a single standard-format master file to replace the 894 original files. ISP converted 817 of the 894 original files before its funding apparently ran out. The ISP-reformatted NURE data files have been released by the USGS on CD-ROM (Lower 48 States, Hoffman and Buttleman, 1994; Alaska, Hoffman and Buttleman, 1996). A description of each NURE database field, derived from a draft NURE HSSR data format manual (unpubl. commun., Stan Moll, ISP, Oct 7, 1988), was included in a readme file on each CD-ROM. That original manual was incomplete and assumed that the reformatting process had gone to completion. A lot of vital information was not included. Efforts to correct that manual and the NURE data revealed a large number of problems and missing data. As a result of the frustrating process of cleaning and re-cleaning data from the ISP-reformatted NURE files, a new NURE HSSR data format was developed. This work represents a totally new attempt to reformat the original NURE files into 2 consistent database structures; one for water samples and a second for sediment samples, on a quadrangle by quadrangle basis, from the original NURE files. Although this USGS-reformatted NURE HSSR data format is different than that created by the ISP, many of their ideas were incorporated and expanded in this effort. All of the data from each quadrangle are being examined thoroughly in an attempt to eliminate problems, to combine partial or duplicate records, to convert all coding to a common scheme, and to identify problems even if they can not be solved at this time.

Open-File Report

Characteristics of the Alaskan 1-Km Advanced Very High Resolution Radiometer data sets used for analysis of vegetation biophysical properties

In this study, data characteristics for composited, multitemporal Advanced Very High Resolution Radiometer data sets for Alaska were assessed for a 7- year period from 1991 to 1997. This involved consideration of the satellite sensors used, data processing performed, and data set compilation, along with an analysis of acquisition date, solar zenith angle, satellite viewing angle, presence of clouds, and registration accuracy for each year. Each year?s worth of data are available on CD-ROM in byte format. All data sets have an initial start date of April 1, but had varying ending dates (mid-September to late October) because of satellite sensor malfunction or the presence of clouds or snow; no data set extended beyond October 31. Satellite scan angles were summarized in seven categories: data obtained at nadir, data within 30, 40, and 55 degrees of nadir, data greater than 55 degrees off nadir, and proportions of the data representing east or west look angles. Minimum, maximum, and average solar zenith angles were provided for each period. Estimates of cloud cover for each period were based on three tests: reflectance gross cloud test, channel 3 minus channel 4, and channel 4 minus channel 5. Registration accuracy was estimated using a gray-level autocorrelation technique. Results of this investigation indicate that the composited data available on CD-ROM should be useful for a number of different regional assessments of Earth cover properties. However, caution is advised when using these data because (1) loss in precision from the conversion to a byte format, (2) low sun angles and high viewing angles in the September and October data, and (3) registration inaccuracies of 2 to 8 pixels.

Alaska

Organic-carbon sequestration in soil/sediment of the Mississippi River deltaic plain — Data; landscape distribution, storage, and inventory; accumulation rates; and recent loss, including a post-Katrina preliminary analysis

Soil/sediment of the Mississippi River deltaic plain (MRDP) in southeastern Louisiana is rich in organic carbon (OC). The MRDP contains about 2 percent of all OC in the surface meter of soil/sediment in the Mississippi River Basin (MRB). Environments within the MRDP differ in soil/sediment organic carbon (SOC) accumulation rate, storage, and inventory. The focus of this study was twofold: (1) develop a database for OC and bulk density for MRDP soil/sediment; and (2) estimate SOC storage, inventory, and accumulation rates for the dominant environments (brackish, intermediate, and fresh marsh; natural levee; distributary; backswamp; and swamp) in the MRDP. Comparative studies were conducted to determine which field and laboratory methods result in the most accurate and reproducible bulk-density values for each marsh environment. Sampling methods included push-core, vibracore, peat borer, and Hargis1 sampler. Bulk-density data for cores taken by the "short push-core method" proved to be more internally consistent than data for samples collected by other methods. Laboratory methods to estimate OC concentration and inorganic-constituent concentration included mass spectrometry, coulometry, and loss-on-ignition. For the sampled MRDP environments, these methods were comparable. SOC storage was calculated for each core with adequate OC and bulk-density data. SOC inventory was calculated using core-specific data from this study and available published and unpublished pedon data linked to SSURGO2 map units. Sample age was estimated using isotopic cesium ( 37 Cs), lead ( 210 Pb), and carbon ( 14 C), elemental Pb, palynomorphs, other stratigraphic markers, and written history. SOC accumulation rates were estimated for each core with adequate age data. Cesium-137 profiles for marsh soil/sediment are the least ambiguous. Levee and distributary 137 Cs profiles show the effects of intermittent allochthonous input and/or sediment resuspension. Cesium-137 and 210 Pb data gave the most consistent and interpretable information for age estimations of soil/sediment deposited during the 1900s. For several cores, isotopic 14 C and 137 Cs data allowed the 1963-64 nuclear weapons testing (NWT) peak-activity datum to be placed within a few-centimeter depth interval. In some cores, a too old 14 C age (when compared to 137 Cs and microstratigraphic-marker data) is the probable result of old carbon bound to clay minerals incorporated into the organic soil/sediment. Elemental Pb coupled with Pb source-function data allowed age estimation for soil/sediment that accumulated during the late 1920s through the 1980s. Exotic pollen (for example, Vigna unguiculata and Alternanthera philoxeroides ) and other microstratigraphic indicators (for example, carbon spherules) allowed age estimations for marsh soil/sediment deposited during the settlement of New Orleans (1717-20) through the early 1900s. For this study, MRDP distributary and swamp environments were each represented by only one core, backswamp environment by two cores, all other environments by three or more cores. MRDP core data for the surface meter soil/sediment indicate that (1) coastal marshes, abandoned distributaries, and swamps have regional SOC-storage values >16 kg m -2 ; (2) swamps and abandoned distributaries have the highest SOC storage values (swamp, 44.8 kg m -2 ; abandoned distributary, 50.9 kg m -2 ); (3) fresh-to-brackish marsh environments have the second highest site-specific SOC-storage values; and (4) site-specific marsh SOC storage values decrease as the salinity of the environment increases (fresh-marsh, 36.2 kg m -2 ; intermediate marsh, 26.2 kg m -2 ; brackish marsh, 21.5 kg m -2 ). This inverse relation between salinity and SOC storage is opposite the regional systematic increase in SOC storage with increasing salinity that is evident when SOC storage is mapped by linking pedon data to SSURGO map units (fresh marsh, 47 kg m -2 ; intermediate marsh, 67 kg m -2 ; brackish marsh, 75 kg m -2 ; and salt marsh, 80 kg m -2 ). MRDP core data for this study also indicate that levees and backswamp have regional SOC-storage values <16 kg m -2 . Group-mean SOC storage for cores from these environments are natural levee (17.0 kg m -2 ) and backswamp (14.1 kg m -2 ). An estimate for the SOC inventory in the surface meter of soil/sediment in the MRDP can be made using the SSURGO mapped portion of the coastal-marsh vegetative-type map (13,236 km 2 , land-only area) published by the Louisiana Department of Wildlife and Fisheries and U.S. Geological Survey (1997). This area has a SOC inventory (surface meter) of 677 Tg (slightly more than 2 percent of the 30,289 Tg SOC inventory for the MRB). The MRDP (6,180 km2, land-only area) has an estimated SOC inventory of 397 Tg. Most of the MRDP is located within the SSURGO mapped coastal marshlands. The entire MRDP, including water, has an area of about 10,800 km 2 . Using the ratio of total MRDP area to SSURGO mapped MRDP area as an adjustment, the MRDP SOC inventory is estimated at 694 Tg. This larger estimate of 694 Tg for the SOC inventory is probably more realistic, because it is reasonable to assume that the marsh sediments overlain by shallow water have comparable SOC storage to that of the adjacent land areas. MRDP core data for this study indicate that there is some variability in long-term SOC mass-accumulation rates for centuries and millennia and that this variability may indicate important geologic changes or changes in land use. However, the consistency of the range in rates of SOC accumulation through time suggests a remarkable degree of marsh sustainability throughout the Holocene, including the recent period of significant marsh modification/channelization for human use. One example of marsh sustainability is its present ability to function as a SOC sink even with Louisiana's large-scale coastal land loss during the last several decades. With coastal-marsh restoration efforts, this sink potential will increase. Looking to the future, a total of 1,101 g m -2 yr -1 SOC is projected to be lost from all of coastal Louisiana (U.S. Army Corps of Engineers, Louisiana Coastal Area (LCA) subprovinces 1-4; not just the MRDP) through coastal erosion from year 2000 to 2050. This translates to a projected SOC-loss rate of about 0.20 percent per year. The recent Hurricanes Katrina and Rita, which devastated the Louisiana coast during late August and late September 2005, transformed about 259 km 2 (100 mi 2 ) of marsh to open water (U.S. Geological Survey, 2005). To the extent that some or all of this land loss is permanent, this result equates to a SOC loss of about 15 Tg. This estimate is based on the year-2000 15,153-km 2 land area for the LCA study area that includes LCA subprovince 4. Using the year-2000 land area, the LCA study area had an estimated SOC inventory of 858 Tg. The estimated 15 Tg SOC loss attributable to Hurricanes Katrina and Rita is 1.7 percent of the year-2000 LCA inventory and 2.3 percent of the year-2000 MRDP inventory. If this SOC loss is included in the projection for the year 2050, then the MRDP would either remain a source with a net SOC loss of 3 Tg or become a weak sink with a net SOC gain of 4 Tg. These estimates are lower bounds for potential SOC flux because they are only for the surface meter of landmass.

Louisiana

Sediment-transport investigations of the upper Yellowstone River, Montana, 1999 through 2001: Data collection, analysis, and simulation of sediment transport

The upper Yellowstone River in Montana is an important State and national water resource, providing recreational, agricultural, and commercial benefits. Floods in 1996 and 1997, with recorded peak discharges having recurrence intervals close to 100 years, caused substantial streambank erosion and hill- slope mass wasting. Large quantities of sand-, gravel-, and cobble-sized material entrained by the flood flows became flood-bar deposits, creating a source of sediment available for transport during future floods. The flood damage and resulting sedimentation raised concerns about potential streambank-stabilization projects and how the river and riparian corridor might be managed in the future. The U.S. Geological Survey, in cooperation with the Park Conservation District, the Montana Department of Transportation, and the U.S. Army Corps of Engineers, investigated sediment transport in the upper Yellowstone River near Livingston from 1999 through 2001 as part of a cumulative effects study to provide a scientific basis for future river management decisions. The purpose of this report is to present the results of data collection, analysis, and simulation of sediment transport for the upper Yellowstone River. The study area included a 13.5-mile study reach of the upper Yellowstone River where substantial sediment transport occurred in 1996 and 1997. In this study area, the upper Yellowstone River is a high gradient, coarse-bed stream having a slope of about 0.0028 foot per foot or more than 14 feet per mile. The study area drains about 3,551 square miles, and runoff results primarily from snowmelt during the spring and summer months. As part of sediment-transport investigations, the U.S. Geological Survey surveyed river cross sections, characterized streambed-material particle size using particle counts and sieve analyses, and collected bedload- and suspended-sediment data during three runoff seasons (1999-2001). Data were collected for stream discharges that ranged from 2,220 cubic feet per second (typical of pre- and post-runoff discharge) to 25,100 cubic feet per second (about 125 percent of bankfull discharge). The distribution of streambed-material particle size was determined, and sediment-transport curves for bedload discharge, suspended-sediment discharge, and total-sediment discharge were developed. The threshold values of streamflow and average stream velocity needed for initiation of bedload transport for selected sediment-size classes showed that little to no bedload was transported for an average stream velocity below about 3 feet per second, and the only particle size transported as bedload at that velocity was sand. Over the range of stream discharges sampled and with silt- and finer-sized particles excluded, bedload discharge averaged about 18 percent of the total-sediment discharge, equal to bedload discharge plus suspended-sediment discharge. At the lowest and highest stream discharges sampled, bedload was, respectively, less than about 2 percent and about 30 percent of the total-sediment discharge. Over the range of stream discharges sampled, the sand-sized part of the total suspended-sediment discharge averaged about 48 percent, where the total suspended-sediment discharge included sand-, silt- and finer-sized particles. At the lowest and highest stream discharges sampled, the sand-sized part of the total suspended-sediment discharge was, respectively, less than about 16 percent and about 50 percent of the total suspended-sediment discharge. The sediment-transport curves were compared to curves for selected sites in the western United States having drainage areas ranging from 21 square miles to over 20,000 square miles. Daily sediment loads transported at bankfull discharge were calculated for each site and results were plotted in relation to drainage area. Results based on the 1999-2001 data-collection period indicate that the estimated daily bedload transported at bankfull discharge in the upper Yellowstone River exceeded the envelope line that bounds the upper end of the data for other selected sites in the Northern Rocky Mountains and is similar in magnitude to that for selected sites in Alaska having braided channels and glacial and snowmelt runoff. Similar comparisons for suspended sediment indicate that daily suspended-sediment load at bankfull discharge is relatively high in the upper Yellowstone River, plotting slightly above the envelope line that bounds the upper end of the data for other selected sites in the Northern Rocky Mountains. Sediment data were used to develop individual transport equations for seven size classes of sediment ranging from small cobbles to very fine sand. A step-wise regression procedure relating sediment discharge to important hydraulic variables showed that average stream velocity was the only significant variable at the 95-percent confidence level. Bedload and suspended-sediment data and equations indicate that more sand is transported for a given velocity than any other particle size, and very little sand-size sediment load is transported below an average stream velocity of about 2.5 feet per second. Transport of coarser-sized sediment (limited to bedload) becomes very little for an average velocity less than about 3.5 feet per second. Results for the 1999-2001 data-collection period indicate that sediment transport in the upper Yellowstone River tends to be limited more by the transport capacity of the stream (capacity or transport limited), than to the availability of sediment in the watershed (supply limited). Sediment data collected and analyzed were used to simulate sediment transport in the study reach using the BRIdge Stream Tube model for Alluvial River Simulation, or BRI-STARS computer model. The model was calibrated and verified using selected data from historical runoff periods. Simulated total-sediment loads, on a reach-averaged basis, were in good agreement with the total-sediment loads determined from the transport curve for the 2-year flood hydrograph but were considerably smaller for the total-sediment loads determined from the transport curve for the 50-, 100-, and 500-year flood hydrographs. The differences probably were largely due to the inability of the model to simulate streambank erosion, hillslope mass-wasting, and other channel-widening processes, which had supplied substantial quantities of sediment to the channel during the 1996 and 1997 floods, and probably continued to contribute to the sediment load in the subsequent years (1999-2001) when the data were collected. Furthermore, the transport curve was applied beyond the measured data for the highest discharges, and may thus be unreliable. Also, the transport curve derived from only limited data may not apply over the full duration of the hydrograph and sediment might be transported over only a portion of the hydrograph, especially for rivers like the upper Yellowstone where snowmelt runoff predominates. The true sediment discharge is, therefore, unknown and might be closer to the simulated values than to the values calculated from the transport curve.

Montana

Quality of pesticide data for groundwater analyzed for the National Water-Quality Assessment Project, 2013–18

The National Water-Quality Assessment (NAWQA) Project of the U.S. Geological Survey (USGS) submitted nearly 1,900 samples collected from groundwater sites across the United States in 2013–18 for analysis of 225 pesticide compounds (pesticides and pesticide degradates, hereafter referred to as “pesticides”) by USGS National Water Quality Laboratory schedule 2437 (S2437). For the associated NAWQA study of pesticide occurrence and concentration in groundwater, and for other studies using pesticide results determined by S2437, it is necessary to assess the ability of reported results to meet data-quality requirements that will allow study objectives to be achieved. This assessment of the quality of S2437 results reported in 2013–18 examined data from field and laboratory quality-control samples, along with third-party performance assessment samples, to estimate bias and variability and to identify their potential sources, with an emphasis on implications for the interpretation of pesticide data for groundwater. Results indicate that measurements produced by the S2437 method for most pesticides have bias and variability that would be considered acceptable for many interpretative studies, which could therefore use the results without qualification or censoring. However, the reported data for a subset of pesticides have the potential for unacceptable contamination bias, high or low recovery bias, or high variability as a consequence of method performance and (or) nonlaboratory factors that could preclude their use for certain common objectives or could necessitate adjustment or qualification to meet those objectives. Based on data for laboratory blanks, censoring of some detections for a subset of pesticides reported by the laboratory in environmental samples might be necessary or desirable to avoid an unacceptably high likelihood of a false-positive result caused by laboratory contamination. The 90-percent upper confidence limit for the 95th percentile of laboratory blank concentration equals or exceeds the minimum reported groundwater concentration in at least 1 water year for 28 pesticides. During at least 1 water year, this upper confidence limit exceeds the maximum laboratory detection limit for 17 pesticides and exceeds the maximum laboratory reporting limit for 3 pesticides (ametryn, atrazine, and diazinon). The level of contamination indicated by this upper confidence limit should not substantially affect the suitability of reported environmental concentrations for any compound for comparison with corresponding human-health benchmarks. Despite being subjected to the same laboratory processes as laboratory blanks, field blanks indicated little evidence of contamination bias. This observation could largely be the consequence of data-reporting practices, which utilize detections in laboratory blanks to censor results in associated field samples (including blanks and environmental samples) when relative concentrations indicate that a result could have a substantial contribution from laboratory contamination. Laboratory censoring appears likely to reduce the risk of false-positive results in environmental samples below the level that laboratory blank results alone would imply. Whereas data available for third-party blind blank samples analyzed in 2018 indicate that only propoxur had any false-positive results, data for pesticides that were not spiked into blind spike samples analyzed in 2013–18 indicate that the false-positive rates for 31 pesticides exceeded 1 percent when considering only detections reported at concentrations greater than the maximum detection limit. Although about half of these pesticides lack substantial supporting evidence of contamination bias based on laboratory blank or field blank detections, indicating that spiking issues or degradation of parent compounds within the spiked samples might be a contributing factor to some false-positive results, these results indicate the need to closely examine detections reported for some pesticides in environmental samples analyzed during a similar period for possible contributions from contamination bias. Data for blind spike samples that were spiked at concentrations above the maximum reporting limit indicate that false-negative rates for eight pesticides exceed 10 percent; substantial low bias could affect results reported for these pesticides in environmental samples analyzed during a similar period. Data for laboratory reagent spikes, which measure recovery of pesticides in blank water, show little evidence for unacceptable recovery bias for S2437 pesticides. However, field matrix spikes, which measure recovery of pesticides in environmental matrices, indicate that degradation and (or) matrix effects could result in moderate to substantial low bias for groundwater results for several pesticides. Low bias could cause some reported concentrations to be categorized as being below a benchmark when the actual concentration in groundwater is greater than the benchmark. Occurrence and concentrations in groundwater could be substantially underrepresented for six pesticides with benchmarks (1H-1,2,4-triazole, asulam, bifenthrin, cis-permethrin, fenbutatin oxide, and naled) that have median recoveries between zero and 50 percent in field matrix spikes. Two compounds (didealkylatrazine and 2-hydroxy-6-ethylamino-4-amino-s-triazine) have median recoveries near or greater than 150 percent in field matrix spikes, indicating a substantial high bias. Plots of data for all spike types show clear changes in the typical recovery with time for some pesticides, which would require further examination for evaluation of temporal trends in environmental concentrations. Data for laboratory reagent spikes indicate that nearly all S2437 pesticides have acceptable variability resulting from random measurement error. Only two compounds (fenbutatin oxide and naled) have F-pseudosigma values greater than 30 percent for recovery, which implies the potential for relatively high variability in reported concentrations and could affect comparison of concentrations to benchmarks and determination of whether concentrations for samples collected at separate locations or times are truly different with a specified level of confidence. Data for third-party blind spike samples show relatively high variability for a greater number of pesticides, although these results likely reflect the influence of degradation and (or) differences in the magnitude and variability of concentrations used for blind spikes relative to laboratory reagent spikes. Detailed analysis of variability using field replicate data is possible for only 12 pesticides on S2437; low variability in analyte detection and concentration is indicated for most of these pesticides in groundwater.

Scientific Investigations Report