Geology Reports⌕ Search

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Identification of hydraulic conductivity structure in sand and gravel aquifers: Cape Cod data set

This study evaluates commonly used geostatistical methods to assess reproduction of hydraulic conductivity (K) structure and sensitivity under limiting amounts of data. Extensive conductivity measurements from the Cape Cod sand and gravel aquifer are used to evaluate two geostatistical estimation methods, conditional mean as an estimate and ordinary kriging, and two stochastic simulation methods, simulated annealing and sequential Gaussian simulation. Our results indicate that for relatively homogeneous sand and gravel aquifers such as the Cape Cod aquifer, neither estimation methods nor stochastic simulation methods give highly accurate point predictions of hydraulic conductivity despite the high density of collected data. Although the stochastic simulation methods yielded higher errors than the estimation methods, the stochastic simulation methods yielded better reproduction of the measured In (K) distribution and better reproduction of local contrasts in In (K). The inability of kriging to reproduce high In (K) values, as reaffirmed by this study, provides a strong instigation for choosing stochastic simulation methods to generate conductivity fields when performing fine-scale contaminant transport modeling. Results also indicate that estimation error is relatively insensitive to the number of hydraulic conductivity measurements so long as more than a threshold number of data are used to condition the realizations. This threshold occurs for the Cape Cod site when there are approximately three conductivity measurements per integral volume. The lack of improvement with additional data suggests that although fine-scale hydraulic conductivity structure is evident in the variogram, it is not accurately reproduced by geostatistical estimation methods. If the Cape Cod aquifer spatial conductivity characteristics are indicative of other sand and gravel deposits, then the results on predictive error versus data collection obtained here have significant practical consequences for site characterization. Heavily sampled sand and gravel aquifers, such as Cape Cod and Borden, may have large amounts of redundant data, while in more common real world settings, our results suggest that denser data collection will likely improve understanding of permeability structure.

Water Resources Research↗

Stress orientation determined from fault slip data in Hampel Wash area, Nevada, and its relation to contemporary regional stress field

Fault-slip data were collected from an area of relatively young faulting in a seismically active part of the Nevada Test Site 12 km NW of Mercury, Nevada. The data come primarily from intensely faulted Miocene tuffaceous sedimentary rocks in Hampel Wash, which is bounded on the north by the Quaternary ENE trending Rock Valley fault and on the south by a parallel unnamed fault. Data from faults with known sense of displacement exhibit a bimodal distribution of slip angles (rakes). Faults exhibiting steep rakes (typically 75° to 90°) cluster about a N30°–35°E strike; most dip 65° to 80°. Faults having shallow rakes (generally less than 20°) exhibit a wide range of strikes (from N6°W to N80°E) and mostly dip between 80° and 90°. The predominant N30°–35°E strike of the steep-rake faults and the quasi-conjugate nature of a consistent subset of the shallowrake faults suggest a maximum horizontal stress orientation of about N30°–35°E and a least horizontal principal stress direction of N55°–60°W. Analysis of the data using a least squares iterative inversion to determine a mean deviatoric principal stress tensor indicates a normal-faulting stress regime (S 1 vertical) with principal stress axes in approximately horizontal and vertical directions (S 1 , trend = N 19°E and plunge = 82°N; S 2 , N30°E and 8°S; and S 3 , N60°W and 2°E). The maximum horizontal stress, S 2 , was found to be nearly intermediate in magnitude between S 1 and S 3 . The N60°W least horizontal principal stress orientation obtained from the fault-slip inversion agrees with our geometric analysis of the data and is consistent with a modern least horizontal principal stress orientation of N50°–70°W inferred from earthquake focal mechanisms, well bore breakouts, and hydraulic fracturing measurements in the vicinity of the Nevada Test Site. This solution fits all the data well, except for a subset of strike-slip faults that strike N30°–45°E, subparallel to the normal faults of the data set. Nearly pure dip-slip and pure strike-slip movement on similarly oriented faults, however, cannot be accommodated in a single stress regime. Superposed sets of striae observed on some faults suggest temporal rotations of the regional stress field or local rotations within the region of the fault zone.

Nevada↗

Best practices for genetic and genomic data archiving

Genetic and genomic data are collected for a vast array of scientific and applied purposes. Despite mandates for public archiving, data are typically used only by the generating authors. The reuse of genetic and genomic datasets remains uncommon because it is difficult, if not impossible, due to non-standard archiving practices and lack of contextual metadata. But as the new field of macrogenetics is demonstrating, if genetic data and their metadata were more accessible and FAIR (findable, accessible, interoperable and reusable) compliant, they could be reused for many additional purposes. We discuss the main challenges with existing genetic and genomic data archives, and suggest best practices for archiving genetic and genomic data. Recognizing that this is a longstanding issue due to little formal data management training within the fields of ecology and evolution, we highlight steps that research institutions and publishers could take to improve data archiving.

Nature, Ecology and Evolution↗

Spatial fuel data products of the LANDFIRE Project

The Landscape Fire and Resource Management Planning Tools (LANDFIRE) Project is mapping wildland fuels, vegetation, and fire regime characteristics across the United States. The LANDFIRE project is unique because of its national scope, creating an integrated product suite at 30-m spatial resolution and complete spatial coverage of all lands within the 50 states. Here we describe development of the LANDFIRE wildland fuels data layers for the conterminous 48 states: surface fire behavior fuel models, canopy bulk density, canopy base height, canopy cover, and canopy height. Surface fire behavior fuel models are mapped by developing crosswalks to vegetation structure and composition created by LANDFIRE. Canopy fuels are mapped using regression trees relating field-referenced estimates of canopy base height and canopy bulk density to satellite imagery, biophysical gradients and vegetation structure and composition data. Here we focus on the methods and data used to create the fuel data products, discuss problems encountered with the data, provide an accuracy assessment, demonstrate recent use of the data during the 2007 fire season, and discuss ideas for updating, maintaining and improving LANDFIRE fuel data products.

International Journal of Wildland Fire↗

Effects of temporal variability in ground data collection on classification accuracy

This research tested whether the timing of ground data collection can significantly impact the accuracy of land cover classification. Ft. Riley Military Reservation, Kansas, USA was used to test this hypothesis. The U.S. Army's Land Condition Trend Analysis (LCTA) data annually collected at military bases was used to ground truth disturbance patterns. Ground data collected over an entire growing season and data collected one year after the imagery had a kappa statistic of 0.33. When using ground data from only within two weeks of image acquisition the kappa statistic improved to 0.55. Potential sources of this discrepancy are identified. These data demonstrate that there can be significant amounts of land cover change within a narrow time window on military reservations. To accurately conduct land cover classification at military reservations, ground data need to be collected in as narrow a window of time as possible and be closely synchronized with the date of the satellite imagery.

Geocarto International↗

Geospatial data mining for digital raster mapping

We performed an in-depth literature survey to identify the most popular data mining approaches that have been applied for raster mapping of ecological parameters through the use of Geographic Information Systems (GIS) and remotely sensed data. Popular data mining approaches included decision trees or “data mining” trees which consist of regression and classification trees, random forests, neural networks, and support vector machines. The advantages of each data mining approach as well as approaches to avoid overfitting are subsequently discussed. We also provide suggestions and examples for the mapping of problematic variables or classes, future or historical projections, and avoidance of model bias. Finally, we address the separate issues of parallel processing, error mapping, and incorporation of “no data” values into modeling processes. Given the improved availability of digital spatial products and remote sensing products, data mining approaches combined with parallel processing potentials should greatly improve the quality and extent of ecological datasets.

GIScience and Remote Sensing↗

Investigating lake-area dynamics across a permafrost-thaw spectrum using airborne electromagnetic surveys and remote sensing time-series data in Yukon Flats, Alaska

Lakes in boreal lowlands cycle carbon and supply an important source of freshwater for wildlife and migratory waterfowl. The abundance and distribution of these lakes are supported, in part, by permafrost distribution, which is subject to change. Relationships between permafrost thaw and lake dynamics remain poorly known in most boreal regions. Here, new airborne electromagnetic (AEM) data collected during June 2010 and February 2016 were used to constrain deep permafrost distribution. AEM data were coupled with Landsat-derived lake surface-area data from 1979 through 2011 to inform temporal lake behavior changes in the 35 500- km 2 Yukon Flats ecoregion of Alaska. Together, over 1500 km of AEM data, and roughly 30 years of Landsat data were used to explore processes that drive lake dynamics across a variety of permafrost thaw states not possible in studies conducted with satellite imagery or field measurements alone. Clustered time-series data identified lakes with similar temporal dynamics. Clusters possessed similarities in lake permanence (i.e. ephemeral versus perennial), subsurface permafrost distribution, and proximity to rivers and streams. Of the clustered lakes, ~66% are inferred to have at least intermittent connectivity with other surface-water features, ~19% are inferred to have shallow subsurface connectivity to other surface water features that served as a low-pass filter for hydroclimatic fluctuations, and ~15% appear to be isolated by surrounding permafrost (i.e. no connectivity). Integrated analysis of AEM and Landsat data reveals a progression from relatively synchronous lake dynamics among disconnected lakes in the most spatially continuous, thick permafrost to quite high spatiotemporal heterogeneity in lake behavior among variably-connected lakes in regions with notably less continuous permafrost. Variability can be explained by the preferential development of thawed permeable gravel pathways for lateral water redistribution in this area. The general spatial progression in permafrost thaw state and lake area behavior may be extended to the temporal dimension. However, extensive permafrost thaw, beyond what is currently observed, is expected to promote ubiquitous subsurface connectivity, eventually evolving to a state of increased lake synchronicity.

Alaska↗

A distributed pipeline for DIDSON data processing

Technological advances in the field of ecology allow data on ecological systems to be collected at high resolution, both temporally and spatially. Devices such as Dual-frequency Identification Sonar (DIDSON) can be deployed in aquatic environments for extended periods and easily generate several terabytes of underwater surveillance data which may need to be processed multiple times. Due to the large amount of data generated and need for flexibility in processing, a distributed pipeline was constructed for DIDSON data making use of the Hadoop ecosystem. The pipeline is capable of ingesting raw DIDSON data, transforming the acoustic data to images, filtering the images, detecting and extracting motion, and generating feature data for machine learning and classification. All of the tasks in the pipeline can be run in parallel and the framework allows for custom processing. Applications of the pipeline include monitoring migration times, determining the presence of a particular species, estimating population size and other fishery management tasks.

Conference Paper↗

Regional land cover characterization using multiple sources of intermediate-scale data

Many United States federal agencies need accurate, intermediate scaled, land cover information. While many techniques and approaches have been successfully used to classify land cover in relatively small regions, there are substantial problems in applying these techniques to large multi-scene regions. An evaluation was conducted of the multiple layer land characteristics data base approach for generating large area land cover information. Mosaicked leaves-on Landsat thematic mapper scenes were used in conjunction with leaves-off thematic mapper data, digital elevation (and derived slope, aspect and shaded relief) data, population census information, defense meteorological satellite program "city lights" data, land use and land cover data, digital line graph data, and national wetlands inventory data to derive land cover information. This approach was evaluated for Region III of the United States Environmental Protection Agency (middle Atlantic states).

Conference Paper↗

Radiometric recalibration procedure for Landsat-5 Thematic Mapper data

The Landsat-5 (L5) satellite was launched on March 01, 1984, with a design life of three years. Incredibly, the L5 Thematic Mapper (TM) has collected data for 23 years. Over this time, the detectors have aged, and its radiometric characteristics have changed since launch. The calibration procedures and parameters have also changed with time. Revised radiometric calibrations have improved the radiometric accuracy of recently processed data; however, users with data that were processed prior to the calibration update do not benefit from the revisions. A procedure has been developed to give users the ability to recalibrate their existing Level 1 (Ll) products without having to purchase reprocessed data from the U.S. Geological Survey (USGS). The accuracy of the recalibration is dependent on the knowledge of the prior calibration applied to the data. The “Work Order” File, included with standard National Land Archive Production System (NLAPS) data products, gives parameters that define the applied calibration. These are the Internal Calibrator (IC) calibration parameters or the default prelaunch calibration, if there were problems with the IC calibration. This paper details the recalibration procedure for data processed using IC, in which users have the Work Order file.

Conference Paper↗

Evaluation of candidate Landsat Data Gap Sensors

The capabilities of the currently operational Landsat satellites may be lost before the launch of the follow-on Landsat Data Continuity Mission (LDCM), thus producing a gap in the Landsat data record and the National Satellite Land Remote Sensing Data Archive (NSLRSDA). In anticipation of a gap, the Federal agencies responsible for Landsat program management, the National Aeronautics and Space Administration (NASA) and the Department of Interior (DOI) U. S. Geological Survey (USGS), convened a Landsat Data Gap Study Team (LDGST). The study team assessed the basic characteristics of multiple systems and identified sensors aboard the China-Brazil Earth Resources Satellite (CBERS-2) and the Indian Remote Sensing (IRS-P6) ResourceSat-1 satellite as the most promising sources of Landsat-like data. The sensors include the combination of CBERS-2 Infrared Multi-spectral Scanner (IRMSS) and High Resolution Charged Coupled Device (CCD), as well as the IRS-P6 Advanced Wide Field Sensor (AWiFS) and the IRS-P6 Linear Imaging Self Scanning Sensor (LISS-III). The study team concluded that more robust technical evaluations of data and sensor performance are required before gap mitigation strategies can be fully formulated. A technical report is made available that summarizes the results from those evaluations, including the initial data characterization and science utility evaluation. The report can be accessed at http://calval.cr.usgs.gov/LDGST.php.

Conference Paper↗

An automated cross-correlation based event detection technique and its application to surface passive data set

In studies on heavy oil, shale reservoirs, tight gas and enhanced geothermal systems, the use of surface passive seismic data to monitor induced microseismicity due to the fluid flow in the subsurface is becoming more common. However, in most studies passive seismic records contain days and months of data and manually analysing the data can be expensive and inaccurate. Moreover, in the presence of noise, detecting the arrival of weak microseismic events becomes challenging. Hence, the use of an automated, accurate and computationally fast technique for event detection in passive seismic data is essential. The conventional automatic event identification algorithm computes a running-window energy ratio of the short-term average to the long-term average of the passive seismic data for each trace. We show that for the common case of a low signal-to-noise ratio in surface passive records, the conventional method is not sufficiently effective at event identification. Here, we extend the conventional algorithm by introducing a technique that is based on the cross-correlation of the energy ratios computed by the conventional method. With our technique we can measure the similarities amongst the computed energy ratios at different traces. Our approach is successful at improving the detectability of events with a low signal-to-noise ratio that are not detectable with the conventional algorithm. Also, our algorithm has the advantage to identify if an event is common to all stations (a regional event) or to a limited number of stations (a local event). We provide examples of applying our technique to synthetic data and a field surface passive data set recorded at a geothermal site.

Geophysical Prospecting↗

Estimating animal resource selection from telemetry data using point process models

Analyses of animal resource selection functions (RSF) using data collected from relocations of individuals via remote telemetry devices have become commonplace. Increasing technological advances, however, have produced statistical challenges in analysing such highly autocorrelated data. Weighted distribution methods have been proposed for analysing RSFs with telemetry data. However, they can be computationally challenging due to an intractable normalizing constant and cannot be aggregated (i.e. collapsed) over time to make space-only inference. In this study, we take a conceptually different approach to modelling animal telemetry data for making RSF inference. We consider the telemetry data to be a realization of a space–time point process. Under the point process paradigm, the times of the relocations are also considered to be random rather than fixed. We show the point process models we propose are a generalization of the weighted distribution telemetry models. By generalizing the weighted model, we can access several numerical techniques for evaluating point process likelihoods that make use of common statistical software. Thus, the analysis methods can be readily implemented by animal ecologists. In addition to ease of computation, the point process models can be aggregated over time by marginalizing over the temporal component of the model. This allows a full range of models to be constructed for RSF analysis at the individual movement level up to the study area level. To demonstrate the analysis of telemetry data with the point process approach, we analysed a data set of telemetry locations from northern fur seals (Callorhinus ursinus) in the Pribilof Islands, Alaska. Both a space–time and an aggregated space-only model were fitted. At the individual level, the space–time analysis showed little selection relative to the habitat covariates. However, at the study area level, the space-only model showed strong selection relative to the covariates.

Alaska↗

Integrating presence-only and detection/non-detection data to estimate distributions and expected abundance of difficult-to-monitor species on a landscape-scale

Estimating species distribution and abundance is foundational to effective management and conservation. Using an integrated species distribution model that combines presence-only data from various sources with detection/non-detection data from structured surveys, we estimated the distribution and expected abundance of three difficult-to-monitor mammals of management concern across New York State, namely, coyotes ( Canis latrans ), bobcats ( Lynx rufus ) and black bears ( Ursus americanus ). Three distinct landscape-scale camera trap surveys provided detection/non-detection data over 9 years between 2013 and 2021, and we augmented those data with incidental records of our focal species from public repositories. We used an inhomogeneous Poisson point process to construct an integrated model that fit both data types simultaneously. We demonstrate a simple application of spatial point density of all species records in the accessed public databases to inform the thinning process to account for unknown spatial sampling in the presence-only data, often referred to as the ‘magic covariate’. Using this approach, we examine habitat associations and provide spatially explicit estimates in expected abundance across the entirety of New York State for all three focal species. As expected, coyotes were the most widely distributed and abundant species, with a strong positive association with agricultural land uses. Bobcats exhibited low expected abundance throughout the state and showed positive associations with deciduous forest and forest edge, and a negative association with road density. Finally, we observed considerable spatial variation in abundance of black bears with expected abundance increasing in association with various forest cover and composition covariates and decreasing with crop cover. We present insights into habitat associations and spatial variation in abundance, and provide management implications for each of the species of interest. Synthesis and applications . Our integrated modelling method allows for managers to use citizen sightings combined with detection/non-detection surveys to estimate robust indices of abundance for both high- and low-density, and wide-spread versus patchily distributed species. Through comparison with previous studies, we highlight how broad-scale programmes, such as the statewide efforts to estimate species distributions undertaken here, can benefit substantively from integrated models that leverage additional data (here, incidental records) from a larger region of space, and thus capture more landscape heterogeneity than is plausible within formalized surveys alone.

New York↗

A new framework for analysing automated acoustic species detection data: Occupancy estimation and optimization of recordings post-processing

The development and use of automated species-detection technologies, such as acoustic recorders, for monitoring wildlife are rapidly expanding. Automated classification algorithms provide a cost- and time-effective means to process information-rich data, but often at the cost of additional detection errors. Appropriate methods are necessary to analyse such data while dealing with the different types of detection errors. We developed a hierarchical modelling framework for estimating species occupancy from automated species-detection data. We explore design and optimization of data post-processing procedures to account for detection errors and generate accurate estimates. Our proposed method accounts for both imperfect detection and false positive errors and utilizes information about both occurrence and abundance of detections to improve estimation. Using simulations, we show that our method provides much more accurate estimates than models ignoring the abundance of detections. The same findings are reached when we apply the methods to two real datasets on North American frogs surveyed with acoustic recorders. When false positives occur, estimator accuracy can be improved when a subset of detections produced by the classification algorithm is post-validated by a human observer. We use simulations to investigate the relationship between accuracy and effort spent on post-validation, and found that very accurate occupancy estimates can be obtained with as little as 1% of data being validated. Automated monitoring of wildlife provides opportunity and challenges. Our methods for analysing automated species-detection data help to meet key challenges unique to these data and will prove useful for many wildlife monitoring programs.

Methods in Ecology and Evolution↗

On the reliability of N‐mixture models for count data

N‐mixture models describe count data replicated in time and across sites in terms of abundance N and detectability p . They are popular because they allow inference about N while controlling for factors that influence p without the need for marking animals. Using a capture–recapture perspective, we show that the loss of information that results from not marking animals is critical, making reliable statistical modeling of N and p problematic using just count data. One cannot reliably fit a model in which the detection probabilities are distinct among repeat visits as this model is overspecified. This makes uncontrolled variation in p problematic. By counter example, we show that even if p is constant after adjusting for covariate effects (the “constant p ” assumption) scientifically plausible alternative models in which N (or its expectation) is non‐identifiable or does not even exist as a parameter, lead to data that are practically indistinguishable from data generated under an N‐mixture model. This is particularly the case for sparse data as is commonly seen in applications. We conclude that under the constant p assumption reliable inference is only possible for relative abundance in the absence of questionable and/or untestable assumptions or with better quality data than seen in typical applications. Relative abundance models for counts can be readily fitted using Poisson regression in standard software such as R and are sufficiently flexible to allow controlling for p through the use covariates while simultaneously modeling variation in relative abundance. If users require estimates of absolute abundance, they should collect auxiliary data that help with estimation of p .

Biometrics↗

Accounting for imperfect detection and survey bias in statistical analysis of presence-only data

Aim During the past decade ecologists have attempted to estimate the parameters of species distribution models by combining locations of species presence observed in opportunistic surveys with spatially referenced covariates of occurrence. Several statistical models have been proposed for the analysis of presence-only data, but these models have largely ignored the effects of imperfect detection and survey bias. In this paper I describe a model-based approach for the analysis of presence-only data that accounts for errors in the detection of individuals and for biased selection of survey locations. Innovation I develop a hierarchical, statistical model that allows presence-only data to be analysed in conjunction with data acquired independently in planned surveys. One component of the model specifies the spatial distribution of individuals within a bounded, geographic region as a realization of a spatial point process. A second component of the model specifies two kinds of observations, the detection of individuals encountered during opportunistic surveys and the detection of individuals encountered during planned surveys. Main conclusions Using mathematical proof and simulation-based comparisons, I demonstrate that biases induced by errors in detection or biased selection of survey locations can be reduced or eliminated by using the hierarchical model to analyse presence-only data in conjunction with counts observed in planned surveys. I show that a relatively small number of high-quality data (from planned surveys) can be used to leverage the information in presence-only observations, which usually have broad spatial coverage but may not be informative of both occurrence and detectability of individuals. Because a variety of sampling protocols can be used in planned surveys, this approach to the analysis of presence-only data is widely applicable. In addition, since the point-process model is formulated at the level of an individual, it can be extended to account for biological interactions between individuals and temporal changes in their spatial distributions.

Global Ecology and Biogeography↗

Combined use of flowmeter and time-drawdown data to estimate hydraulic conductivities in layered aquifer systems

The vertical distribution of hydraulic conductivity in layered aquifer systems commonly is needed for model simulations of ground-water flow and transport. In previous studies, time-drawdown data or flowmeter data were used individually, but not in combination, to estimate hydraulic conductivity. In this study, flowmeter data and time-drawdown data collected from a long-screened production well and nearby monitoring wells are combined to estimate the vertical distribution of hydraulic conductivity in a complex multilayer coastal aquifer system. Flowmeter measurements recorded as a function of depth delineate nonuniform inflow to the wellbore, and this information is used to better discretize the vertical distribution of hydraulic conductivity using analytical and numerical methods. The time-drawdown data complement the flowmeter data by giving insight into the hydraulic response of aquitards when flow rates within the wellbore are below the detection limit of the flowmeter. The combination of these field data allows for the testing of alternative conceptual models of radial flow to the wellbore.

Ground Water↗