Geology Reports⌕ Search

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

An introduction and practical guide to use of the Soil-Vegetation Inventory Method (SVIM) data

Long-term vegetation dynamics across public rangelands in the western United States are not well understood because of the lack of large-scale, readily available historic datasets. The Bureau of Land Management’s Soil-Vegetation Inventory Method (SVIM) program was implemented between 1977 and 1983 across 14 western states, but the data have not been easily accessible. We introduce the SVIM vegetation cover dataset in a georeferenced, digital format; summarize how the data were collected; and discuss potential limitations and biases. We demonstrate how SVIM data can be compared with contemporary monitoring datasets to quantify changes in vegetation associated with wildfire and the abundance of exotic invasive species. Specifically, we compare SVIM vegetation cover data with cover data collected by BLM’s Assessment, Inventory, and Monitoring (AIM) program (2011–2016) in a focal area in the northern Great Basin. We address issues associated with analyzing and interpreting data from these distinct programs, including differences in survey methods and potential biases introduced by spatial and temporal variation in sampling. We compared SVIM and AIM survey methods at 44 plots and found that percent cover estimates had high correspondence for all measured functional groups. Comparisons between historic SVIM data and recent AIM data documented significant declines in the occupancy and cover of native shrubs and native perennial forbs , and a significant increase in exotic annual forbs. Wildfire was a driver of change for some functional groups, with greater change occurring in AIM plots that burned between the two time periods compared with those that did not. Our results are consistent with previous studies showing that many native shrub-dominated plant communities in the Great Basin have been replaced by exotic annuals. Our study demonstrates that SVIM data will be an important resource for researchers interested in quantifying vegetation change through time across public rangelands in the western United States.

Rangeland Ecology and Management↗

Optimal interpolation analysis of leaf area index using MODIS data

A simple data analysis technique for vegetation leaf area index (LAI) using Moderate Resolution Imaging Spectroradiometer (MODIS) data is presented. The objective is to generate LAI data that is appropriate for numerical weather prediction. A series of techniques and procedures which includes data quality control, time-series data smoothing, and simple data analysis is applied. The LAI analysis is an optimal combination of the MODIS observations and derived climatology, depending on their associated errors σ o and σ c . The “best estimate” LAI is derived from a simple three-point smoothing technique combined with a selection of maximum LAI (after data quality control) values to ensure a higher quality. The LAI climatology is a time smoothed mean value of the “best estimate” LAI during the years of 2002–2004. The observation error is obtained by comparing the MODIS observed LAI with the “best estimate” of the LAI, and the climatological error is obtained by comparing the “best estimate” of LAI with the climatological LAI value. The LAI analysis is the result of a weighting between these two errors. Demonstration of the method described in this paper is presented for the 15-km grid of Meteorological Service of Canada (MSC)'s regional version of the numerical weather prediction model. The final LAI analyses have a relatively smooth temporal evolution, which makes them more appropriate for environmental prediction than the original MODIS LAI observation data. They are also more realistic than the LAI data currently used operationally at the MSC which is based on land-cover databases.

Remote Sensing of Environment↗

PaCTS 1.0: A crowdsourced reporting standard for paleoclimate data

The progress of science is tied to the standardization of measurements, instruments, and data. This is especially true in the Big Data age, where analyzing large data volumes critically hinges on that data being standardized. Accordingly, the lack of community-sanctioned data standards in paleoclimatology has largely precluded the benefits of Big Data advances in the field. Building upon recent efforts to standardize the format and terminology of paleoclimate data, this article describes the Paleoclimate Community reporTing Standard (PaCTS), a crowdsourced reporting standard for such data. PaCTS captures which information should be included when reporting paleoclimate data, with the goal of maximizing the re-use value of paleoclimate datasets, particularly for synthesis work and comparison to climate model simulations. Initiated by the LinkedEarth project, this standard elicitation process involved an international workshop in 2016, various forms of digital community engagement over the next few years, and grassroots working groups. Participants in this process identified important properties across paleoclimate archives, in addition to the reporting of uncertainties and chronologies; they also identified archive-specific properties and distinguished reporting standards for new vs. legacy datasets. This work shows that at least 135 respondents overwhelmingly support a drastic increase in the amount of metadata accompanying paleoclimate datasets. Since such goals are at odds with present practices, we discuss a transparent path towards implementing or revising these recommendations in the near future, using both bottom-up and top-down approaches.

Paleoceanography and Paleoclimatology↗

Toward a new framework to evaluate process-based model configurations and quantify data worth prior to calibration

Model criticism, discrimination, and selection methods often rely on calibrated model outputs. Because calibration can be computationally expensive, model criticism can first be undertaken by assessing model outputs obtained from limited prior parameter ensembles. However, such prior-based methods are often heuristic and do not formalize the notion of balancing model consistency with data and model complexity (i.e., model adequacy). We present a new framework to discriminate among candidate models prior to calibration that formalizes prior-to-calibration model adequacy into a metric to implicitly balance prior model output data coverage with model complexity represented by prior output (co)variance. The prior model adequacy metric “Mahalanobis distance deviation” quantifies the deviation of (a) the set of squared Mahalanobis distances of data from a prior model output distribution from (b) the set of squared Mahalanobis distances of data from their own distribution. A new data worth metric “discernment value” is also presented which quantifies the value of data for screening less-adequate models prior to calibration. Discernment value is calculated from the change in variance of a weighted average of prior model outputs from all candidate models due to less-adequate model outputs receiving lower weight. The framework is demonstrated using a one-dimensional groundwater flow model with eight possible configurations. A synthetic data network is used to test the framework. Results show the framework identifies the candidate models most similar to the true model used to create the synthetic data. Discernment values show variation in the value of different data types and locations for screening less-adequate models.

Water Resources Research↗

Adjustment of total suspended solids data for use in sediment studies

The U.S. Environmental Protection Agency identifies fluvial sediment as the single most widespread pollutant in the Nation's rivers and streams, affecting aquatic habitat, drinking water treatment processes, and recreational uses of rivers, lakes, and estuaries. A significant amount of suspended-sediment data has been produced using the total suspended solids (TSS) laboratory analysis method. An evaluation of data collected and analyzed by the U.S. Geological Survey and others has shown that the variation in TSS analytical results is considerably larger than that for traditional suspended-sediment concentration analyses (SSC) and that the TSS data show a negative bias when compared to SSC data. This paper presents the initial results of a continuing investigation into the differences between TSS and SSC results. It explores possible relations between these differences and other hydrologic data collected at the same stations. A general equation was developed to relate TSS data to SSC data. However, this general equation is not applicable for data from individual stations. Based on these analyses, there appears to be no simple, straightforward way to relate TSS and SSC data unless pairs of TSS and SSC results are available for a station.

Conference Paper↗

North American landscape characterization project: The production of a continental scale three-decade Landsat data set

The North American Landscape Characterization (NALC) project is a component of the National Aeronautics and Space Administration (NASA) Landsat Pathfinder program. Pathfinder projects are focused on the investigation of global change utilizing current remote sensing technologies. The NALC project is a cooperative effort between the U. S. Environmental Protection Agency (EPA), the U.S. Geological Survey (USGS), and NASA to make Landsat data available to the widest possible user community for scientific research and general public interest. The NALC project is principally funded by the EPA Office of Research and Development and the USGS's Earth Resources Observation Systems (EROS) Data Center (EDC). The objectives of the NALC project are to produce standardized remote sensing data sets, develop standardized analysis methods, and derive standardized land cover change products for a large portion of the North American continent (the conterminous United States and Mexico) (Lunetta and Sturdevant, 1993). The standard product is the NALC “triplicate”;, consisting of co‐registered Landsat multispectral scanner data for the years 1973, 1986, and 1991 (plus or minus one year), plus co‐registered 3 arcsecond digital terrain elevation data. Processing began with the 1986 scene, which was precision corrected (with full terrain correction) to a 60 meter Universal Transverse Mercator base. Automated cross‐correlation procedures were used to co‐register the 1970's and 1990's data to the 1980's base, and independent verifications of registration quality were performed on all triplicate components. The pertinent metadata were compiled in a relational database, which includes WRS2 path/rows, scene ID's, image dates, solar azimuth and elevation, verification RMSE's, and the number of verification control points. NALC triplicate data sets are being used for a number of applications, including the analysis of urbanization patterns, dynamics of climatic fluctuations, deforestation studies, and vegetation classification and mapping. These data are being distributed through the Earth Observing System Data and Information System (EOSDIS) Information Management System (IMS) at a cost of $15(U.S.) for each triplicate.

Geocarto International↗

Subsurface structure across the Tacoma Basin, Washington State, using trans-dimensional Bayesian inversion of fundamental mode spatial autocorrelation data

Spatial autocorrelation (SPAC), the azimuthal average of the normalized cross-correlation between equidistant station pairs deployed in a 2-D array, is widely used to image the subsurface structure. However, the rigorous estimate of subsurface structure and its uncertainties as a function of depth using SPAC data is challenging due to the nonlinear relation between the SPAC data and Earth structure as well as the trade-off between depth and velocity. Additionally, data noise is strongly correlated due to data processing (e.g. filtering, stacking from multiple time segments and azimuthal averaging). Most studies do not account for the correlated noise and fix the ratio of compressional-wave velocity ( V P ) to shear-wave velocity ( V s ) (i.e. V P / V s ratio) and the number of layers, both of which are typically unknown. To address these challenges, we develop a hierarchical trans-dimensional Bayesian inversion of fundamental mode of SPAC data that properly accounts for the correlated data noise, samples the V P / V s ratio and relaxes the number of layers (i.e. model parametrization) to be unknown in the inversion. We further examine the limitation of using only fundamental modes in the inversion. Our synthetic experiments show that the inversion recovers an incorrect model unless we sample the correlated noise and V P / V s ratio in the inversion. The inversion is then applied to SPAC data acquired at 19 sites across the Tacoma basin in Washington State to characterize the V s and the time-averaged V s over 30-m depth ( V s 30 ). Our results show that the V s 30 varies from ∼200 to 800 m s −1 . The V s 30 within the basin is higher in the middle and lower on the east and west sides. We find that these V s 30 values vary with geologic unit. The uncertainties for V s 30 are within 20 m s −1 in average except for the most eastern site TB28. Additionally, the uncertainties are greater for deeper depths beneath most of the sites as the sensitivity decreases as a function of depth. The Vs structure as a function of depth is also complex beneath some sites, possibly because the SPAC curves are affected by higher order Rayleigh modes that are not considered in the inversion. To better constrain the deeper V s structure, V s 30 and/or other average measures of V s over depth, additional constraints from complementary data, such as ellipticity or geologic data are needed. Moreover, our synthetic experiments show that higher order modes can have significant effect in the inversion results, particularly when there is a low-velocity layer.

Washington↗

Uncertainty in biological monitoring: a framework for data collection and analysis to account for multiple sources of sampling bias

Biological monitoring programmes are increasingly relying upon large volumes of citizen-science data to improve the scope and spatial coverage of information, challenging the scientific community to develop design and model-based approaches to improve inference. Recent statistical models in ecology have been developed to accommodate false-negative errors, although current work points to false-positive errors as equally important sources of bias. This is of particular concern for the success of any monitoring programme given that rates as small as 3% could lead to the overestimation of the occurrence of rare events by as much as 50%, and even small false-positive rates can severely bias estimates of occurrence dynamics. We present an integrated, computationally efficient Bayesian hierarchical model to correct for false-positive and false-negative errors in detection/non-detection data. Our model combines independent, auxiliary data sources with field observations to improve the estimation of false-positive rates, when a subset of field observations cannot be validated a posteriori or assumed as perfect. We evaluated the performance of the model across a range of occurrence rates, false-positive and false-negative errors, and quantity of auxiliary data. The model performed well under all simulated scenarios, and we were able to identify critical auxiliary data characteristics which resulted in improved inference. We applied our false-positive model to a large-scale, citizen-science monitoring programme for anurans in the north-eastern United States, using auxiliary data from an experiment designed to estimate false-positive error rates. Not correcting for false-positive rates resulted in biased estimates of occupancy in 4 of the 10 anuran species we analysed, leading to an overestimation of the average number of occupied survey routes by as much as 70%. The framework we present for data collection and analysis is able to efficiently provide reliable inference for occurrence patterns using data from a citizen-science monitoring programme. However, our approach is applicable to data generated by any type of research and monitoring programme, independent of skill level or scale, when effort is placed on obtaining auxiliary information on false-positive rates.

Methods in Ecology and Evolution↗

Occupancy models for citizen-science data

Large‐scale citizen‐science projects, such as atlases of species distribution, are an important source of data for macroecological research, for understanding the effects of climate change and other drivers on biodiversity, and for more applied conservation tasks, such as early‐warning systems for biodiversity loss. However, citizen‐science data are challenging to analyse because the observation process has to be taken into account. Typically, the observation process leads to heterogeneous and non‐random sampling, false absences, false detections, and spatial correlations in the data. Increasingly, occupancy models are being used to analyse atlas data. We advocate a dual approach to strengthen inference from citizen science data for the questions the programme is intended to address: (a) the survey design should be chosen with a particular set of questions and associated analysis strategy in mind and (b) the statistical methods should be tailored not only to those questions but also to the specific characteristics of the data. We review the consequences of particular survey design choices that typically need to be made in atlas‐style citizen‐science projects. These include spatial resolution of the sampling units, allocation of effort in space, and collection of information about the observation process. On the analysis side, we review extensions of the basic occupancy models that are frequently necessary with atlas data, including methods for dealing with heterogeneity, non‐independent detections, false detections, and violation of the closure assumption. New technologies, such as cell‐phone apps and fixed remote detection devices, are revolutionizing citizen‐science projects. There is an opportunity to maximize the usefulness of the resulting datasets if the protocols are rooted in robust statistical designs and data analysis issues are being considered. Our review provides guidelines for designing new projects and an overview of the current methods that can be used to analyse data from such projects.

Methods in Ecology and Evolution↗

A genetic algorithm to reduce stream channel cross section data

A genetic algorithm (GA) was used to reduce cross section data for a hypothetical example consisting of 41 data points and for 10 cross sections on the Kootenai River. The number of data points for the Kootenai River cross sections ranged from about 500 to more than 2,500. The GA was applied to reduce the number of data points to a manageable dataset because most models and other software require fewer than 100 data points for management, manipulation, and analysis. Results indicated that the program successfully reduced the data. Fitness values from the genetic algorithm were lower (better) than those in a previous study that used standard procedures of reducing the cross section data. On average, fitnesses were 29 percent lower, and several were about 50 percent lower. Results also showed that cross sections produced by the genetic algorithm were representative of the original section and that near-optimal results could be obtained in a single run, even for large problems. Other data also can be reduced in a method similar to that for cross section data.

Journal of the American Water Resources Associatio↗

An effective noise-suppression technique for surface microseismic data

The presence of strong surface-wave noise in surface microseismic data may decrease the utility of these data. We implement a technique, based on the distinct characteristics that microseismic signal and noise show in the τ‐p domain, to suppress surface-wave noise in microseismic data. Because most microseismic source mechanisms are deviatoric, preprocessing is necessary to correct for the nonuniform radiation pattern prior to transforming the data to the τ‐p domain. We employ a scanning approach, similar to semblance analysis, to test all possible double-couple orientations to determine an estimated orientation that best accounts for the polarity pattern of any microseismic events. We then correct the polarity of the data traces according to this pattern, prior to conducting signal-noise separation in the τ‐p domain. We apply our noise-suppression technique to two surface passive-seismic data sets from different acquisition surveys. The first data set includes a synthetic microseismic event added to field passive noise recorded by an areal receiver array distributed over a Barnett Formation reservoir undergoing hydraulic fracturing. The second data set is field microseismic data recorded by receivers arranged in a star-shaped array, over a Bakken Shale reservoir during a hydraulic-fracturing process. Our technique significantly improves the signal-to-noise ratios of the microseismic events and preserves the waveforms at the individual traces. We illustrate that the enhancement in signal-to-noise ratio also results in improved imaging of the microseismic hypocenter.

Geophysics↗

Multielevation calibration of frequency-domain electromagnetic data

Systematic calibration errors must be taken into account because they can substantially impact the accuracy of inverted subsurface resistivity models derived from frequency-domain electromagnetic data, resulting in potentially misleading interpretations. We have developed an approach that uses data acquired at multiple elevations over the same location to assess calibration errors. A significant advantage is that this method does not require prior knowledge of subsurface properties from borehole or ground geophysical data (though these can be readily incorporated if available), and is, therefore, well suited to remote areas. The multielevation data were used to solve for calibration parameters and a single subsurface resistivity model that are self consistent over all elevations. The deterministic and Bayesian formulations of the multielevation approach illustrate parameter sensitivity and uncertainty using synthetic- and field-data examples. Multiplicative calibration errors (gain and phase) were found to be better resolved at high frequencies and when data were acquired over a relatively conductive area, whereas additive errors (bias) were reasonably resolved over conductive and resistive areas at all frequencies. The Bayesian approach outperformed the deterministic approach when estimating calibration parameters using multielevation data at a single location; however, joint analysis of multielevation data at multiple locations using the deterministic algorithm yielded the most accurate estimates of calibration parameters. Inversion results using calibration-corrected data revealed marked improvement in misfit, lending added confidence to the interpretation of these models.

Geophysics↗

Cumulative uncertainty in measured streamflow and water quality data for small watersheds

The scientific community has not established an adequate understanding of the uncertainty inherent in measured water quality data, which is introduced by four procedural categories: streamflow measurement, sample collection, sample preservation/storage, and laboratory analysis. Although previous research has produced valuable information on relative differences in procedures within these categories, little information is available that compares the procedural categories or presents the cumulative uncertainty in resulting water quality data. As a result, quality control emphasis is often misdirected, and data uncertainty is typically either ignored or accounted for with an arbitrary margin of safety. Faced with the need for scientifically defensible estimates of data uncertainty to support water resource management, the objectives of this research were to: (1) compile selected published information on uncertainty related to measured streamflow and water quality data for small watersheds, (2) use a root mean square error propagation method to compare the uncertainty introduced by each procedural category, and (3) use the error propagation method to determine the cumulative probable uncertainty in measured streamflow, sediment, and nutrient data. Best case, typical, and worst case data quality scenarios were examined. Averaged across all constituents, the calculated cumulative probable uncertainty (%) contributed under typical scenarios ranged from 6% to 19% for streamflow measurement, from 4% to 48% for sample collection, from 2% to 16% for sample preservation/storage, and from 5% to 21% for laboratory analysis. Under typical conditions, errors in storm loads ranged from 8% to 104% for dissolved nutrients, from 8% to 110% for total N and P, and from 7% to 53% for TSS. Results indicated that uncertainty can increase substantially under poor measurement conditions and limited quality control effort. This research provides introductory scientific estimates of uncertainty in measured water quality data. The results and procedures presented should also assist modelers in quantifying the quality of calibration and evaluation data sets, determining model accuracy goals, and evaluating model performance.

Transactions of the ASABE↗

Annual survival of Snail Kites in Florida: Radio telemetry versus capture-resighting data

We estimated annual survival of Snail Kites ( Rostrhamus sociabilis ) in Florida using the Kaplan-Meier estimator with data from 271 radio-tagged birds over a three-year period and capture-recapture (resighting) models with data from 1,319 banded birds over a six-year period. We tested the hypothesis that survival differed among three age classes using both data sources. We tested additional hypotheses about spatial and temporal variation using a combination of data from radio telemetry and single- and multistrata capture-recapture models. Results from these data sets were similar in their indications of the sources of variation in survival, but they differed in some parameter estimates. Both data sources indicated that survival was higher for adults than for juveniles, but they did not support delineation of a subadult age class. Our data also indicated that survival differed among years and regions for juveniles but not for adults. Estimates of juvenile survival using radio telemetry data were higher than estimates using capture-recapture models for two of three years (1992 and 1993). Ancillary evidence based on censored birds indicated that some mortality of radio-tagged juveniles went undetected during those years, resulting in biased estimates. Thus, we have greater confidence in our estimates of juvenile survival using capture-recapture models. Precision of estimates reflected the number of parameters estimated and was surprisingly similar between radio telemetry and single-stratum capture-recapture models, given the substantial differences in sample sizes. Not having to estimate resighting probability likely offsets, to some degree, the smaller sample sizes from our radio telemetry data. Precision of capture-recapture models was lower using multistrata models where region-specific parameters were estimated than using single-stratum models, where spatial variation in parameters was not taken into account.

The Auk↗

Availability Of Ground-Water Data For California, Water Year 2000

The Water Resources Division of the U.S. Geological Survey, in cooperation with Federal, State, and local water agencies, obtains a large amount of data pertaining to the ground-water resources of California each water year. These data constitute a valuable database for developing an improved understanding of the water resources of the State. Beginning with the 1985 water year and continuing through 1993, these data were published in a report series entitled ?Water Resources Data for California, Volume 5. Ground-Water Data.? Prior to the introduction of this series, historical ground-water information was published in U.S. Geological Survey Water-Supply Papers. In 1994, the Volume 5 Ground-Water Data report was discontinued, but data continue to be available in our databases. This Fact Sheet serves as an index to ground-water data for water year 2000. The 2-page report contains a map of California showing the number of wells (by county) with available water-level and water-quality data for water year 2000 (fig. 2) and instructions for obtaining this and other ground-water information contained in the databases of the Water Resources Division, California District.

Fact Sheet↗

User's Manual for the New England Water-Use Data System (NEWUDS)

Water is used in a variety of ways that need to be understood for effective management of water resources. Water-use activities need to be categorized and included in a database management system to understand current water uses and to provide information to water-resource management policy decisionmakers. The New England Water-Use Data System (NEWUDS) is a complex database developed to store water-use information that allows water to be tracked from a point of water-use activity (called a 'Site'), such as withdrawal from a resource (reservoir or aquifer), to a second Site, such as distribution to a user (business or irrigator). NEWUDS conceptual model consists of 10 core entities: system, owner, address, location, site, data source, resource, conveyance, transaction/rate, and alias, with tables available to store user-defined details. Three components--site (with both a From Site and a To Site), a conveyance that connects them, and a transaction/rate associated with the movement of water over a specific time interval form the core of the basic NEWUDS network model. The most important step in correctly translating real-world water-use activities into a storable format in NEWUDS depends on choosing the appropriate sites and linking them correctly in a network to model the flow of water from the initial From Site to the final To Site. Ten water-use networks representing real-world activities are described--three withdrawal networks, three return networks, two user networks, two complex community-system networks. Ten case studies of water use, one for each network, also are included in this manual to illustrate how to compile, store, and retrieve the appropriate data. The sequence of data entry into tables is critical because there are many foreign keys. The recommended core entity sequence is (1) system, (2) owner, (3) address, (4) location, (5) site, (6) data source, (7) resource, (8) conveyance, (9) transaction, and (10) rate; with (11) alias and (12) user-defined detail subject areas populated as needed. After each step in data entry, quality-assurance queries should be run to ensure the data are correctly entered so that it can be retrieved accurately. The point of data storage is retrieval. Several retrieval queries that focus on retrieving only relevant data to specific questions are presented in this manual as examples for the NEWUDS user.

Open-File Report↗

Data for Quaternary faults in western Montana

The "World Map of Major Active Faults" Task Group is compiling published fault data, developing a digital database of the fault data, and preparing a series of maps for the United States and other countries in the western Hemisphere. The data is intended to portray the locations, ages, and activity rates of major earthquake-related features such as faults, folds, and liquefaction features that have geologic evidence of Quaternary (1.6 Ma) deformation. The Western Hemisphere effort is sponsored by International Lithosphere Program (ILP) Task Group II-2; the data compilation, database, and map for the United States is funded largely by the National Earthquake Hazard Reduction Program (NEHRP) through the U.S. Geological Survey. The ILP effort in the Western Hemisphere is coordinated by Michael N. Machette, the digital database is designed and managed Kathleen M. Haller, and map data are digitized and manipulated by Richard L. Dart. In addition to meeting the goals of the Task Group II-2, this effort represents a key contribution to the new Global Seismic Hazards Assessment Program (ILP Task Group II-0) for the International Decade for Natural Disaster Reduction. This compilation, which documents the published data on Quaternary surface faulting in western Montana, is one of many similar state or regional compilations that are planned for the project. Compilations for Arizona (Pearthree, 1998 #2945), Colorado (Widmann and others, 1998 #3441), New Mexico (Machette and others, 1998), and West Texas (Collins and others, 1996 #993) are currently available and the compilation for features east of the Rocky Mountain front will be available in early 2000 (Crone and Wheeler, in press). All are primarily a catalog of data that includes a variety of geographic, geologic, and paleoseismologic parameters for known or assumed Quaternary faults. These data compilations, the digital database, and the companion maps summarize the published information on known tectonic features and present the information in an internally consistent format. The compilations will be available in digital database format on the WorldWide Web in the near future, which will greatly improve their utility. Release of data for individual states and regions within the United States in this text-based format was necessary because of the time required to develop the national database.

Montana↗

Data Model and Relational Database Design for Highway Runoff Water-Quality Metadata

A National highway and urban runoff waterquality metadatabase was developed by the U.S. Geological Survey in cooperation with the Federal Highway Administration as part of the National Highway Runoff Water-Quality Data and Methodology Synthesis (NDAMS). The database was designed to catalog available literature and to document results of the synthesis in a format that would facilitate current and future research on highway and urban runoff. This report documents the design and implementation of the NDAMS relational database, which was designed to provide a catalog of available information and the results of an assessment of the available data. All the citations and the metadata collected during the review process are presented in a stratified metadatabase that contains citations for relevant publications, abstracts (or previa), and reportreview metadata for a sample of selected reports that document results of runoff quality investigations. The database is referred to as a metadatabase because it contains information about available data sets rather than a record of the original data. The database contains the metadata needed to evaluate and characterize how valid, current, complete, comparable, and technically defensible published and available information may be when evaluated for application to the different dataquality objectives as defined by decision makers. This database is a relational database, in that all information is ultimately linked to a given citation in the catalog of available reports. The main database file contains 86 tables consisting of 29 data tables, 11 association tables, and 46 domain tables. The data tables all link to a particular citation, and each data table is focused on one aspect of the information collected in the literature search and the evaluation of available information. This database is implemented in the Microsoft (MS) Access database software because it is widely used within and outside of government and is familiar to many existing and potential customers. The stratified metadatabase design for the NDAMS program is presented in the MS Access file DBDESIGN.mdb and documented with a data dictionary in the NDAMS_DD.mdb file recorded on the CD-ROM. The data dictionary file includes complete documentation of the table names, table descriptions, and information about each of the 419 fields in the database.

Open-File Report↗