Procedures for petroleum resource assessment used by the U.S. Geological Survey; statistical and probabilistic methodology
Explore the source record for details and available documents.
SEARCH · Geology Reports
Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The U.S. Geological Survey has developed a methodology for statistically relating nutrient sources and land-surface characteristics to nutrient loads of streams. The methodology is referred to as SPAtially Referenced Regressions On Watershed attributes (SPARROW), and relates measured stream nutrient loads to nutrient sources using nonlinear statistical regression models. A spatially detailed digital hydrologic network of stream reaches, stream-reach characteristics such as mean streamflow, water velocity, reach length, and travel time, and their associated watersheds supports the regression models. This network serves as the primary framework for spatially referencing potential nutrient source information such as atmospheric deposition, septic systems, point-sources, land use, land cover, and agricultural sources and land-surface characteristics such as land use, land cover, average-annual precipitation and temperature, slope, and soil permeability. In the Chesapeake Bay watershed that covers parts of Delaware, Maryland, Pennsylvania, New York, Virginia, West Virginia, and Washington D.C., SPARROW was used to generate models estimating loads of total nitrogen and total phosphorus representing 1987 and 1992 land-surface conditions. The 1987 models used a hydrologic network derived from an enhanced version of the U.S. Environmental Protection Agency's digital River Reach File, and course resolution Digital Elevation Models (DEMs). A new hydrologic network was created to support the 1992 models by generating stream reaches representing surface-water pathways defined by flow direction and flow accumulation algorithms from higher resolution DEMs. On a reach-by-reach basis, stream reach characteristics essential to the modeling were transferred to the newly generated pathways or reaches from the enhanced River Reach File used to support the 1987 models. To complete the new network, watersheds for each reach were generated using the direction of surface-water flow derived from the DEMs. This network improves upon existing digital stream data by increasing the level of spatial detail and providing consistency between the reach locations and topography. The hydrologic network also aids in illustrating the spatial patterns of predicted nutrient loads and sources contributed locally to each stream, and the percentages of nutrient load that reach Chesapeake Bay.
We present a methodology for statistical analysis of randomly located marine sediment point data, and apply it to the US continental shelf portions of usSEABED mean grain size records. The usSEABED database, like many modern, large environmental datasets, is heterogeneous and interdisciplinary. We statistically test the database as a source of mean grain size data, and from it provide a first examination of regional seafloor sediment variability across the entire US continental shelf. Data derived from laboratory analyses ("extracted") and from word-based descriptions ("parsed") are treated separately, and they are compared statistically and deterministically. Data records are selected for spatial analysis by their location within sample regions: polygonal areas defined in ArcGIS chosen by geography, water depth, and data sufficiency. We derive isotropic, binned semivariograms from the data, and invert these for estimates of noise variance, field variance, and decorrelation distance. The highly erratic nature of the semivariograms is a result both of the random locations of the data and of the high level of data uncertainty (noise). This decorrelates the data covariance matrix for the inversion, and largely prevents robust estimation of the fractal dimension. Our comparison of the extracted and parsed mean grain size data demonstrates important differences between the two. In particular, extracted measurements generally produce finer mean grain sizes, lower noise variance, and lower field variance than parsed values. Such relationships can be used to derive a regionally dependent conversion factor between the two. Our analysis of sample regions on the US continental shelf revealed considerable geographic variability in the estimated statistical parameters of field variance and decorrelation distance. Some regional relationships are evident, and overall there is a tendency for field variance to be higher where the average mean grain size is finer grained. Surprisingly, parsed and extracted noise magnitudes correlate with each other, which may indicate that some portion of the data variability that we identify as "noise" is caused by real grain size variability at very short scales. Our analyses demonstrate that by applying a bias-correction proxy, usSEABED data can be used to generate reliable interpolated maps of regional mean grain size and sediment character.
Determining appropriate statistical distributions for modeling animal count data is important for accurate estimation of abundance, distribution, and trends. In the case of sea ducks along the U.S. Atlantic coast, managers want to estimate local and regional abundance to detect and track population declines, to define areas of high and low use, and to predict the impact of future habitat change on populations. In this paper, we used a modified marked point process to model survey data that recorded flock sizes of Common eiders, Long-tailed ducks, and Black, Surf, and White-winged scoters. The data come from an experimental aerial survey, conducted by the United States Fish & Wildlife Service (USFWS) Division of Migratory Bird Management, during which east-west transects were flown along the Atlantic Coast from Maine to Florida during the winters of 2009–2011. To model the number of flocks per transect (the points), we compared the fit of four statistical distributions (zero-inflated Poisson, zero-inflated geometric, zero-inflated negative binomial and negative binomial) to data on the number of species-specific sea duck flocks that were recorded for each transect flown. To model the flock sizes (the marks), we compared the fit of flock size data for each species to seven statistical distributions: positive Poisson, positive negative binomial, positive geometric, logarithmic, discretized lognormal, zeta and Yule–Simon. Akaike’s Information Criterion and Vuong’s closeness tests indicated that the negative binomial and discretized lognormal were the best distributions for all species for the points and marks, respectively. These findings have important implications for estimating sea duck abundances as the discretized lognormal is a more skewed distribution than the Poisson and negative binomial, which are frequently used to model avian counts; the lognormal is also less heavy-tailed than the power law distributions (e.g., zeta and Yule–Simon), which are becoming increasingly popular for group size modeling. Choosing appropriate statistical distributions for modeling flock size data is fundamental to accurately estimating population summaries, determining required survey effort, and assessing and propagating uncertainty through decision-making processes.
A common population characteristic of interest in animal ecology studies pertains to the selection of resources. That is, given the resources available to animals, what do they ultimately choose to use? A variety of statistical approaches have been employed to examine this question and each has advantages and disadvantages with respect to the form of available data and the properties of estimators given model assumptions. A wealth of high resolution telemetry data are now being collected to study animal population movement and space use and these data present both challenges and opportunities for statistical inference. We summarize traditional methods for resource selection and then describe several extensions to deal with measurement uncertainty and an explicit movement process that exists in studies involving high-resolution telemetry data. Our approach uses a correlated random walk movement model to obtain temporally varying use and availability distributions that are employed in a weighted distribution context to estimate selection coefficients. The temporally varying coefficients are then weighted by their contribution to selection and combined to provide inference at the population level. The result is an intuitive and accessible statistical procedure that uses readily available software and is computationally feasible for large datasets. These methods are demonstrated using data collected as part of a large-scale mountain lion monitoring study in Colorado, USA.
Accessing and effectively sampling the off-channel habitats that are considered crucial for early life stages of freshwater fishes constitute a difficult challenge when common ichthyoplankton survey methods, such as push nets, are used. We describe a new method of deploying push nets from jet-propelled kayaks to enable the sampling of previously inaccessible off-channel habitats. The described rig is also functional in more open and accessible habitats, such as the main channel of rivers or reservoirs. Although further evaluation is necessary to ensure that results are comparable across studies, the described push-net system offers a statistically rigorous methodology that generates replicate samples from a wide range of freshwater habitats that were previously inaccessible to this gear type.
In their analysis of the U.S. Geological Survey’s (USGS) “Did You Feel It?” (DYFI) data Hough and Martin (2021) claim, among other assertions, that the following: Socioeconomic and geopolitical factors can introduce biases in the USGS’ characterization of earthquakes and their effects, especially if online data collection systems are not designed to be broadly accessible; These biases can, in turn, potentially cascade in myriad ways, potentially shaping our understanding of an earthquake’s impact and the characterization of seismic hazard; and Caution should be urged when relying on data from the DYFI system to characterize the distribution of shaking from large earthquakes in India and other parts of the world (outside of the United States). Claims of inequity in access, systematic data biases, or urging caution in the usage of data from critical governmental earthquake information systems should not be made, nor taken, lightly. Several assertions made by Hough and Martin (hereafter, H&M) about the nature of DYFI contributors—and the data they provide—leave a false narrative concerning DYFI system accessibility and quality that H&M have not adequately substantiated. I describe several shortcomings of H&M’s demographic statistics and methodology, focusing on four main concerns. First, DYFI has revolutionized and greatly facilitated access to reporting intensities, in contrast to H&M claims to the contrary. Second, because DYFI does not directly collect demographic data other than the observer’s location, any demographic analyses require extraordinary inferences, well outside the normal bounds of sociodemographic analyses. Third, independent of accessibility and the geographic distribution of contributions from the public, the macroseismic data collected are nonetheless representative of the shaking and impact at each location, of quality, rapid, and thus extremely useful. Lastly, H&M fail to cite critical and pertinent prior, highly relevant scholarly studies, and as such, they misrepresent the novelty of their own work as well as miss key practical matters detailed in those prior studies. Prior to rebutting what H&M claim DYFI does not do, I will remind the reader the ways in which DYFI excels.
Point counts are a standard sampling procedure for many bird species, but lingering concerns still exist about the quality of information produced from the method. It is well known that variation in observer ability and environmental conditions can influence the detection probability of birds in point counts, but many biologists have been reluctant to abandon point counts in favor of more intensive approaches to counting. However, over the past few years a variety of statistical and methodological developments have begun to provide practical ways of overcoming some of the problems with point counts. We describe some of these approaches, and show how they can be integrated into standard point count protocols to greatly enhance the quality of the information. Several tools now exist for estimation of detection probability of birds during counts, including distance sampling, double observer methods, time-depletion (removal) methods, and hybrid methods that combine these approaches. Many counts are conducted in habitats that make auditory detection of birds much more likely than visual detection. As a framework for understanding detection probability during such counts, we propose separating two components of the probability a bird is detected during a count into (1) the probability a bird vocalizes during the count and (2) the probability this vocalization is detected by an observer. In addition, we propose that some measure of the area sampled during a count is necessary for valid inferences about bird populations. This can be done by employing fixed-radius counts or more sophisticated distance-sampling models. We recommend any studies employing point counts be designed to estimate detection probability and to include a measure of the area sampled.
A statistical forecast methodology exploits large-scale patterns in monthly U.S. Climatological Division Palmer Drought Severity Index (PDSI) values over a wide region and several seasons to predict area burned in western U.S. wildfires by ecosystem province a season in advance. The forecast model, which is based on canonical correlations, indicates that a few characteristic patterns determine predicted wildfire season area burned. Strong negative associations between anomalous soil moisture (inferred from PDSI) immediately prior to the fire season and area burned dominate in most higher elevation forested provinces, while strong positive associations between anomalous soil moisture a year prior to the fire season and area burned dominate in desert and shrub and grassland provinces. In much of the western U.S., above- and below-normal fire season forecasts were successful 57% of the time or better, as compared with a 33% skill for a random guess, and with a low probability of being surprised by a fire season at the opposite extreme of that forecast.
Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.
Shorebirds Charadrii are prime candidates for population decline because of their dependence on wetlands that are being lost at a rapid pace. Thirty-six of the 49 species of shorebirds that breed in North America spend most of the year in Latin America. Because populations of most species breed and winter at remote sites, it may be most feasible to monitor their numbers at migration stopovers. In this study, we used statistical trend analysis methods, developed for the North America Breeding Bird Survey, to analyze data on shorebird populations during southbound migration in the United States. Survey data were collected by volunteers in the International Shorebird Survey (ISS). The analyses indicate that whimbrels Numenius phaeopus , short-billed dowitchers Limnodromus griseus , and sanderlings Calidris alba have undergone statistically significant declines. Methodological concerns over both the ISS and the trend analysis procedures are discussed in detail and biological interpretations of the results are suggested.
Climate change is altering wildfire behavior and vegetation regimes in California’s forested ecosystems. Present day fires are seeing an increase in high burn severity area and high severity patch size. The ability to predict future burn severity patterns would support better policy and land management decisions. Here we demonstrate a methodology to first, statistically estimate individual burn severity classes at 30 meters and second, cluster and smooth high severity patches onto a landscape. Our goal here was not to exactly replicate observed burn severity maps, but rather to utilize observed maps as one realization of a random process dependent on climate, topography, fire weather, and fuels, to inform creation of additional realizations through our simulation technique. We developed two sets of empirical models with two different vegetation datasets to test if coarse vegetation could accurately model for burn severity. While visual acuity can be used to assess the performance of our simulation process, we also employ the Ripley’s K function to compare spatial point processes at different scales to test if the simulation is capturing an appropriate amount of clustering. We utilize FRAGSTATS to obtain high severity patch metrics to test the contiguity of our high severity simulation. Ripley’s K function helped identify the number of clustering iterations and FRAGSTATS showed how different focal window sizes affected our ability to cluster high severity patches. High severity patch simulations was comparable between the coarse and fine resolution vegetation datasets. Improving our ability to simulate burn severity will help advance our understanding of the potential influence of land and fuels management on ecosystem-level response variables that are important for decision-makers. Simulated burn severity maps can support managing habitat and estimating risks of habitat loss, protecting infrastructure and homes, improving future wildfire emissions projections, and better mapping and planning for fuels treatment scenarios.
Characterization of wave climate by bulk wave parameters is insufficient for many coastal studies, including those focused on assessing coastal hazards and long-term wave climate influences on coastal evolution. This issue is particularly relevant for studies using statistical downscaling of atmospheric fields to local wave conditions, which are often multimodal in large ocean basins (e.g. the Pacific). Swell may be generated in vastly different wave generation regions, yielding complex wave spectra that are inadequately represented by a single set of bulk wave parameters. Furthermore, the relationship between atmospheric systems and local wave conditions is complicated by variations in arrival time of wave groups from different parts of the basin. Here, we address these two challenges by improving upon the spatiotemporal definition of the atmospheric predictor used in statistical downscaling of local wave climate. The improved methodology separates the local wave spectrum into “wave families,” defined by spectral peaks and discrete generation regions, and relates atmospheric conditions in distant regions of the ocean basin to local wave conditions by incorporating travel times computed from effective energy flux across the ocean basin. When applied to locations with multimodal wave spectra, including Southern California and Trujillo, Peru, the new methodology improves the ability of the statistical model to project significant wave height, peak period, and direction for each wave family, retaining more information from the full wave spectrum. This work is the base of statistical downscaling by weather types, which has recently been applied to coastal flooding and morphodynamic applications.
Current land-management decisions that affect the persistence of native salmonids are often influenced by studies of individual sites that are selected based on judgment and convenience. Although this approach is useful for some purposes, extrapolating results to areas that were not sampled is statistically inappropriate because the sampling design is usually biased. Therefore, in recent investigations of coastal cutthroat trout (Oncorhynchus clarki clarki) located above natural barriers to anadromous salmonids, we used a methodology for extending the statistical scope of inference. The purpose of this paper is to apply geospatial tools to identify a population of watersheds and develop a probability-based sampling design for coastal cutthroat trout in western Oregon, USA. The population of mid-size watersheds (500-5800 ha) west of the Cascade Range divide was derived from watershed delineations based on digital elevation models. Because a database with locations of isolated populations of coastal cutthroat trout did not exist, a sampling frame of isolated watersheds containing cutthroat trout had to be developed. After the sampling frame of watersheds was established, isolated watersheds with coastal cutthroat trout were stratified by ecoregion and erosion potential based on dominant bedrock lithology (i.e., sedimentary and igneous). A stratified random sample of 60 watersheds was selected with proportional allocation in each stratum. By comparing watershed drainage areas of streams in the general population to those in the sampling frame and the resulting sample (n = 60), we were able to evaluate the how representative the subset of watersheds was in relation to the population of watersheds. Geospatial tools provided a relatively inexpensive means to generate the information necessary to develop a statistically robust, probability-based sampling design.
This document provides a methodology for extracting grain statistics from 8-bit color and grayscale images of thin sections of glacier ice—a subset of physical properties measurements typically performed on ice cores. This type of analysis is most commonly used to characterize the evolution of ice-crystal size, shape, and intercrystalline spatial relations within a large body of ice sampled by deep ice-coring projects from which paleoclimate records will be developed. However, such information is equally useful for investigating the stress state and physical responses of ice to stresses within a glacier. The methods of analysis presented here go hand-in-hand with the analysis of ice fabrics (aggregate crystal orientations) and, when combined with fabric analysis, provide a powerful method for investigating the dynamic recrystallization and deformation behaviors of bodies of ice in motion. The procedures described in this document compose a step-by-step handbook for a specific image acquisition and data reduction system built in support of U.S. Geological Survey ice analysis projects, but the general methodology can be used with any combination of image processing and analysis software. The specific approaches in this document use the FoveaPro 4 plug-in toolset to Adobe Photoshop CS5 Extended but it can be carried out equally well, though somewhat less conveniently, with software such as the image processing toolbox in MATLAB, Image-Pro Plus, or ImageJ.
Estimation of lean mass and lipid levels in birds involves the derivation of predictive equations that relate morphological measurements and, more recently, total body electrical conductivity (TOBEC) indices to known lean and lipid masses. Using cross-validation techniques, we evaluated the ability of several published and new predictive equations to estimate lean and lipid mass of Semipalmated Sandpipers (Calidris pusilla) and White-rumped Sandpipers ( C. fuscicollis ). We also tested ideas of Morton et al. (1991), who stated that current statistical approaches to TOBEC methodology misrepresent precision in estimating body fat. Three published interspecific equations using TOBEC indices predicted lean and lipid masses of our sample of birds with average errors of 8-28% and 53-155%, respectively. A new two-species equation relating lean mass and TOBEC indices revealed average errors of 4.6% and 23.2% in predicting lean and lipid mass, respectively. New intraspecific equations that estimate lipid mass directly from body mass, morphological measurements, and TOBEC indices yielded about a 13% error in lipid estimates. Body mass and morphological measurements explained a substantial portion of the variance (about 90%) in fat mass of both species. Addition of TOBEC indices improved the predictive model more for the smaller than for the larger sandpiper. TOBEC indices explained an additional 7.8% and 2.6% of the variance in fat mass and reduced the minimum breadth of prediction intervals by 0.95 g (32%) and 0.39 g (13%) for Semipalmated and White-rumped Sandpipers, respectively. The breadth of prediction intervals for models used to predict fat levels of individual birds must be considered when interpreting the resultant lipid estimates.
The 2007 Energy Independence and Security Act (Public Law 110–140) directs the U.S. Geological Survey (USGS) to conduct a national assessment of potential geologic storage resources for carbon dioxide (CO 2 ) and to consult with other Federal and State agencies to locate the pertinent geological data needed for the assessment. The geologic sequestration of CO 2 is one possible way to mitigate its effects on climate change. The methodology used for the national CO 2 assessment (Open-File Report 2010-1127; http://pubs.usgs.gov/of/2010/1127/) is based on previous USGS probabilistic oil and gas assessment methodologies. The methodology is non-economic and intended to be used at regional to subbasinal scales. The operational unit of the assessment is a storage assessment unit (SAU), composed of a porous storage formation with fluid flow and an overlying sealing unit with low permeability. Assessments are conducted at the SAU level and are aggregated to basinal and regional results. This report identifies and contains geologic descriptions of SAUs in separate packages of sedimentary rocks within the assessed basin and focuses on the particular characteristics, specified in the methodology, that influence the potential CO 2 storage resource in those SAUs. Specific descriptions of the SAU boundaries as well as their sealing and reservoir units are included. Properties for each SAU such as depth to top, gross thickness, net porous thickness, porosity, permeability, groundwater quality, and structural reservoir traps are provided to illustrate geologic factors critical to the assessment. Although assessment results are not contained in this report, the geologic information included here will be employed, as specified in the methodology, to calculate a statistical Monte Carlo-based distribution of potential storage space in the various SAUs. Figures in this report show SAU boundaries and cell maps of well penetrations through the sealing unit into the top of the storage formation. Wells sharing the same well borehole are treated as a single penetration. Cell maps show the number of penetrating wells within one square mile and are derived from interpretations of incompletely attributed well data, a digital compilation that is known not to include all drilling. The USGS does not expect to know the location of all wells and cannot guarantee the amount of drilling through specific formations in any given cell shown on cell maps.
The 2007 Energy Independence and Security Act (Public Law 110-140) directs the U.S. Geological Survey (USGS) to conduct a national assessment of potential geologic storage resources for carbon dioxide (CO 2 ). The methodology used for the national CO 2 assessment is non-economic and intended to be used at regional to subbasinal scales. This report identifies and contains geologic descriptions of twelve storage assessment units (SAUs) in six separate packages of sedimentary rock within the Hanna, Laramie, and Shirley Basins of Wyoming. It focuses on the particular characteristics, specified in the methodology, that influence the potential CO 2 storage resource in those SAUs. Specific descriptions of SAU boundaries as well as their sealing and reservoir units are included. Properties for each SAU, such as depth to top, gross thickness, net porous thickness, porosity, permeability, groundwater quality, and structural reservoir traps are provided to illustrate geologic factors critical to the assessment. Although assessment results are not contained in this report, the geologic information included herein will be employed, as specified in the methodology, to calculate a statistical Monte Carlo-based distribution of potential storage space in the various SAUs. Figures in this report show SAU boundaries and cell maps of well penetrations through the sealing unit into the top of the storage formation. Cell maps show the number of penetrating wells within one square mile and are derived from interpretations of incompletely attributed well data in a digital compilation that is known not to include all drilling. The USGS does not expect to know the location of all wells and cannot guarantee the amount of drilling through specific formations in any given cell shown on cell maps.