Geology ReportsSearch

SEARCH · Geology Reports

Results for “Journal of Statistical Computation and Simulation”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Efficient estimators for adaptive stratified sequential sampling

In stratified sampling, methods for the allocation of effort among strata usually rely on some measure of within-stratum variance. If we do not have enough information about these variances, adaptive allocation can be used. In adaptive allocation designs, surveys are conducted in two phases. Information from the first phase is used to allocate the remaining units among the strata in the second phase. Brown et al . [ Adaptive two-stage sequential sampling , Popul. Ecol. 50 (2008), pp. 239–245] introduced an adaptive allocation sampling design – where the final sample size was random – and an unbiased estimator. Here, we derive an unbiased variance estimator for the design, and consider a related design where the final sample size is fixed. Having a fixed final sample size can make survey-planning easier. We introduce a biased Horvitz–Thompson type estimator and a biased sample mean type estimator for the sampling designs. We conduct two simulation studies on honey producers in Kurdistan and synthetic zirconium distribution in a region on the moon. Results show that the introduced estimators are more efficient than the available estimators for both variable and fixed sample size designs, and the conventional unbiased estimator of stratified simple random sampling design. In order to evaluate efficiencies of the introduced designs and their estimator furthermore, we first review some well-known adaptive allocation designs and compare their estimator with the introduced estimators. Simulation results show that the introduced estimators are more efficient than available estimators of these well-known adaptive allocation designs.

Journal of Statistical Computation and Simulation

An evaluation of the Bayesian approach to fitting the N-mixture model for use with pseudo-replicated count data

The N-mixture model proposed by Royle in 2004 may be used to approximate the abundance and detection probability of animal species in a given region. In 2006, Royle and Dorazio discussed the advantages of using a Bayesian approach in modelling animal abundance and occurrence using a hierarchical N-mixture model. N-mixture models assume replication on sampling sites, an assumption that may be violated when the site is not closed to changes in abundance during the survey period or when nominal replicates are defined spatially. In this paper, we studied the robustness of a Bayesian approach to fitting the N-mixture model for pseudo-replicated count data. Our simulation results showed that the Bayesian estimates for abundance and detection probability are slightly biased when the actual detection probability is small and are sensitive to the presence of extra variability within local sites.

Journal of Statistical Computation and Simulation

A case study of green tree frog population size estimation by repeated capture-mark-recapture method with individual tagging: A parametric bootstrap method vs. Jolly-Seber method

This paper deals with estimation of a green tree frog population in an urban setting using repeated capture–mark–recapture (CMR) method over several weeks with an individual tagging system which gives rise to a complicated generalization of the hypergeometric distribution. Based on the maximum likelihood estimation, a parametric bootstrap approach is adopted to obtain interval estimates of the weekly population size which is the main objective of our work. The method is computation-based; and programming intensive to implement the algorithm for re-sampling. This method can be applied to estimate the population size of any species based on repeated CMR method at multiple time points. Further, it has been pointed out that the well-known Jolly–Seber method, which is based on some strong assumptions, produces either unrealistic estimates, or may have situations where its assumptions are not valid for our observed data set.

Journal of Statistical Computation and Simulation

Interstation correlation of peak-flow estimates

Equations are given for using the correlation coefficient between annual peaks at a pair of stream gaging stations to estimate the interstation correlation coefficient of estimated T -year peaks, computed standard deviations, and computed skew coefficients at the same pair of stations. The equations are based on statistics computed from samples of annual peaks simulated from normal distributions with built-in interstation correlation. The interstation correlation of 50- year peaks was found to be about equal to the square of the interstation correlation coefficient of the annual peaks.

Journal of Research of the U.S. Geological Survey

Computationally efficient statistical differential equation modeling using homogenization

Statistical models using partial differential equations (PDEs) to describe dynamically evolving natural systems are appearing in the scientific literature with some regularity in recent years. Often such studies seek to characterize the dynamics of temporal or spatio-temporal phenomena such as invasive species, consumer-resource interactions, community evolution, and resource selection. Specifically, in the spatial setting, data are often available at varying spatial and temporal scales. Additionally, the necessary numerical integration of a PDE may be computationally infeasible over the spatial support of interest. We present an approach to impose computationally advantageous changes of support in statistical implementations of PDE models and demonstrate its utility through simulation using a form of PDE known as “ecological diffusion.” We also apply a statistical ecological diffusion model to a data set involving the spread of mountain pine beetle (Dendroctonus ponderosae) in Idaho, USA.

Journal of Agricultural, Biological, and Environme

Cross-fade sampling: Extremely efficient Bayesian inversion for a variety of geophysical problems

This paper introduces cross-fade sampling, a computationally efficient Markov Chain Monte Carlo simulation method that uses a semi-analytical approach to quickly solve Bayesian inverse problems that do not themselves have an analytical solution. Cross-fading is efficient in two ways. First, it requires fewer samples to obtain the same quality simulation of the target probability density function (PDF). Secondly, it is much faster to evaluate the posterior probability of each sample than conventional sampling methods for simulating Bayesian posterior PDFs. Conventional methods require evaluating the prior probability (which describes your a priori constraints) and data likelihood (which describes the fit between the observations and the predictions of the model) for each sample model. However, cross-fading does not require evaluating the data likelihood, meaning that ‘big data’ can be fit with zero additional computational cost. Further, the cross-fading approach can be used to calculate the marginal likelihood associated with a model design, facilitating model comparison and Bayesian model averaging. Topics covered in this paper include derivation of the cross-fade approach and how it can be used to simulate Bayesian posterior PDFs and compute the marginal likelihood, discussion of the class of problems to which cross-fading can be applied (with examples from earthquake statistics, earthquake ground motion modelling, volcanic eruption forecasting, and finite fault slip modelling), demonstration of efficiency relative to existing sampling methods and discussion of how cross-fading can be used to account for prediction errors (i.e. epistemic errors) as part of the geophysical inverse problem.

Geophysical Journal International

A random forest approach for bounded outcome variables

Random forests have become an established tool for classication and regres- sion, in particular in high-dimensional settings and in the presence of non-additive predictor-response relationships. For bounded outcome variables restricted to the unit interval, however, classical modeling approaches based on mean squared error loss may severely suer as they do not account for heteroscedasticity in the data. To address this issue, we propose a random forest approach for relating a beta dis- tributed outcome to a set of explanatory variables. Our approach explicitly makes use of the likelihood function of the beta distribution for the selection of splits dur- ing the tree-building procedure. In each iteration of the tree-building algorithm it chooses one explanatory variable in combination with a split point that maximizes the log-likelihood function of the beta distribution with the parameter estimates de- rived from the nodes of the currently built tree. Results of several simulation studies and an application using data from the U.S.A. National Lakes Assessment Survey demonstrate the properties and usefulness of the method, in particular when com- pared to random forest approaches based on mean squared error loss and parametric regression models.

Journal of Computational and Graphical Statistics

Estimation of population size using open capture-recapture models

One of the most important needs for wildlife managers is an accurate estimate of population size. Yet, for many species, including most marine species and large mammals, accurate and precise estimation of numbers is one of the most difficult of all research challenges. Open-population capture-recapture models have proven useful in many situations to estimate survival probabilities but typically have not been used to estimate population size. We show that open-population models can be used to estimate population size by developing a Horvitz-Thompson-type estimate of population size and an estimator of its variance. Our population size estimate keys on the probability of capture at each trap occasion and therefore is quite general and can be made a function of external covariates measured during the study. Here we define the estimator and investigate its bias, variance, and variance estimator via computer simulation. Computer simulations make extensive use of real data taken from a study of polar bears (Ursus maritimus) in the Beaufort Sea. The population size estimator is shown to be useful because it was negligibly biased in all situations studied. The variance estimator is shown to be useful in all situations, but caution is warranted in cases of extreme capture heterogeneity.

Journal of Agricultural, Biological, and Environme

Assessing first-order emulator inference for physical parameters in nonlinear mechanistic models

We present an approach for estimating physical parameters in nonlinear models that relies on an approximation to the mechanistic model itself for computational efficiency. The proposed methodology is validated and applied in two different modeling scenarios: (a) Simulation and (b) lower trophic level ocean ecosystem model. The approach we develop relies on the ability to predict right singular vectors (resulting from a decomposition of computer model experimental output) based on the computer model input and an experimental set of parameters. Critically, we model the right singular vectors in terms of the model parameters via a nonlinear statistical model. Specifically, we focus our attention on first-order models of these right singular vectors rather than the second-order (covariance) structure.

Journal of Agricultural, Biological, and Environme

Estimator selection for closed-population capture: recapture

For valid statistical inference, it is important to select an appropriate statistical model. In the analysis of capture-recapture data under the closed-population models of Otis et al. (1978), information theoretic and hypothesis testing approaches to model selection are not practical, because some of the models have likelihoods with nonidenti- fiable parameters. A further problem is that, for some of the Otis et al. models, multiple estimators exist but there is no objective basis for deciding which estimator to use for a particular dataset. In CAPTURE, a computer program for estimating parameters un- der the closed models of Otis et al., a linear discriminant classifier is used to select an appropriate model. This classifier frequently selects the incorrect generating model in simulation studies, and it provides no guidance on which estimator to use once a model has been selected. In this study, we develop new classifiers for selecting the best esti- mator (as opposed to the generating model) and evaluate their performance. In addition, we investigate an estimator averaging approach to estimation that is a modification of the model averaging approach described by Buckland et al. (1997). We found that, in general, the overall performance of the new classifiers was unimpressive. In contrast, the estimator averaging approach we investigated performed well.

Journal of Agricultural, Biological, and Environme

An approach to modeling abundance of marine wildlife over space and time using unstructured aerial surveys

Estimating spatial and temporal patterns in abundance is often a goal of ecological studies and can be useful for informing management decisions, such as determining the optimal placement of wildlife protection zones. However, estimating abundance can be difficult in practice, especially over large areas, because of imperfect detection, where individuals are present but not detected because of either availability or observer error. Several methods for estimating abundance that account for imperfect detection exist but can be logistically challenging to implement. We present a simpler approach to some of the more commonly used techniques for estimating the abundance of marine wildlife over space and time from unstructured aerial surveys. This approach combines a spatial model for count data with auxiliary information on detection probability obtained from small-scale or previous studies. We employ generalized linear models and generalized additive models with spatial habitat covariates to illustrate this approach using maximum-likelihood with free, open-source statistical software. This framework is intended to be accessible and flexible, requiring lower survey costs and less computation time than other alternatives for estimating abundance. Indeed, our simulation results show that this approach can reduce computation times, while appropriately characterizing uncertainty, compared to a Bayesian approach. We also present R code for our approach using an example of estimating Florida manatee ( Trichechus manatus latirostris ) abundance in Indian River County in Florida, USA. This approach could be applied to other study systems and marine wildlife species using unstructured aerial surveys.

Florida

Investigations of potential bias in the estimation of lambda using Pradel's (1996) model for capture-recapture data

Pradel's (1996) temporal symmetry model permitting direct estimation and modelling of population growth rate, u i , provides a potentially useful tool for the study of population dynamics using marked animals. Because of its recent publication date, the approach has not seen much use, and there have been virtually no investigations directed at robustness of the resulting estimators. Here we consider several potential sources of bias, all motivated by specific uses of this estimation approach. We consider sampling situations in which the study area expands with time and present an analytic expression for the bias in u i We next consider trap response in capture probabilities and heterogeneous capture probabilities and compute large-sample and simulation-based approximations of resulting bias in u i . These approximations indicate that trap response is an especially important assumption violation that can produce substantial bias. Finally, we consider losses on capture and emphasize the importance of selecting the estimator for u i that is appropriate to the question being addressed. For studies based on only sighting and resighting data, Pradel's (1996) u i ' is the appropriate estimator.

Journal of Applied Statistics

The Bayesian group lasso for confounded spatial data

Generalized linear mixed models for spatial processes are widely used in applied statistics. In many applications of the spatial generalized linear mixed model (SGLMM), the goal is to obtain inference about regression coefficients while achieving optimal predictive ability. When implementing the SGLMM, multicollinearity among covariates and the spatial random effects can make computation challenging and influence inference. We present a Bayesian group lasso prior with a single tuning parameter that can be chosen to optimize predictive ability of the SGLMM and jointly regularize the regression coefficients and spatial random effect. We implement the group lasso SGLMM using efficient Markov chain Monte Carlo (MCMC) algorithms and demonstrate how multicollinearity among covariates and the spatial random effect can be monitored as a derived quantity. To test our method, we compared several parameterizations of the SGLMM using simulated data and two examples from plant ecology and disease ecology. In all examples, problematic levels multicollinearity occurred and influenced sampling efficiency and inference. We found that the group lasso prior resulted in roughly twice the effective sample size for MCMC samples of regression coefficients and can have higher and less variable predictive accuracy based on out-of-sample data when compared to the standard SGLMM.

Journal of Agricultural, Biological, and Environme

Warming is driving decreases in snow fractions while runoff efficiency remains mostly unchanged in snow-covered areas of the western United States

Winter snowfall and accumulation is an important component of the surface water supply in the western United States. In these areas, increasing winter temperatures T associated with global warming can influence the amount of winter precipitation P that falls as snow S . In this study we examine long-term trends in the fraction of winter P that falls as S (Sfrac) for 175 hydrologic units (HUs) in snow-covered areas of the western United States for the period 1951–2014. Because S is a substantial contributor to runoff R across most of the western United States, we also examine long-term trends in water-year runoff efficiency [computed as water-year R /water-year P (Reff)] for the same 175 HUs. In that most S records are short in length, we use model-simulated S and R from a monthly water balance model. Results for Sfrac indicate long-term negative trends for most of the 175 HUs, with negative trends for 139 (~79%) of the HUs being statistically significant at a 95% confidence level ( p = 0.05). Additionally, results indicate that the long-term negative trends in Sfrac have been largely driven by increases in T . In contrast, time series of Reff for the 175 HUs indicate a mix of positive and negative long-term trends, with few trends being statistically significant (at p = 0.05). Although there has been a notable shift in the timing of R to earlier in the year for most HUs, there have not been substantial decreases in water-year R for the 175 HUs.

Arizona, California, Colorado, Idaho, Montana, Nev

A global search inversion for earthquake kinematic rupture history: Application to the 2000 western Tottori, Japan earthquake

[1] We present a two-stage nonlinear technique to invert strong motions records and geodetic data to retrieve the rupture history of an earthquake on a finite fault. To account for the actual rupture complexity, the fault parameters are spatially variable peak slip velocity, slip direction, rupture time and risetime. The unknown parameters are given at the nodes of the subfaults, whereas the parameters within a subfault are allowed to vary through a bilinear interpolation of the nodal values. The forward modeling is performed with a discrete wave number technique, whose Green's functions include the complete response of the vertically varying Earth structure. During the first stage, an algorithm based on the heat-bath simulated annealing generates an ensemble of models that efficiently sample the good data-fitting regions of parameter space. In the second stage (appraisal), the algorithm performs a statistical analysis of the model ensemble and computes a weighted mean model and its standard deviation. This technique, rather than simply looking at the best model, extracts the most stable features of the earthquake rupture that are consistent with the data and gives an estimate of the variability of each model parameter. We present some synthetic tests to show the effectiveness of the method and its robustness to uncertainty of the adopted crustal model. Finally, we apply this inverse technique to the well recorded 2000 western Tottori, Japan, earthquake ( Mw 6.6); we confirm that the rupture process is characterized by large slip (3-4 m) at very shallow depths but, differently from previous studies, we imaged a new slip patch (2-2.5 m) located deeper, between 14 and 18 km depth.

Tottori

Imputation approaches for animal movement modeling

The analysis of telemetry data is common in animal ecological studies. While the collection of telemetry data for individual animals has improved dramatically, the methods to properly account for inherent uncertainties (e.g., measurement error, dependence, barriers to movement) have lagged behind. Still, many new statistical approaches have been developed to infer unknown quantities affecting animal movement or predict movement based on telemetry data. Hierarchical statistical models are useful to account for some of the aforementioned uncertainties, as well as provide population-level inference, but they often come with an increased computational burden. For certain types of statistical models, it is straightforward to provide inference if the latent true animal trajectory is known, but challenging otherwise. In these cases, approaches related to multiple imputation have been employed to account for the uncertainty associated with our knowledge of the latent trajectory. Despite the increasing use of imputation approaches for modeling animal movement, the general sensitivity and accuracy of these methods have not been explored in detail. We provide an introduction to animal movement modeling and describe how imputation approaches may be helpful for certain types of models. We also assess the performance of imputation approaches in two simulation studies. Our simulation studies suggests that inference for model parameters directly related to the location of an individual may be more accurate than inference for parameters associated with higher-order processes such as velocity or acceleration. Finally, we apply these methods to analyze a telemetry data set involving northern fur seals ( Callorhinus ursinus ) in the Bering Sea. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme

Spatial variability in seasonal snowpack trends across the Rio Grande headwaters (1984 - 2017)

This study evaluated the spatial variability of trends in simulated snowpack properties across the Rio Grande headwaters of Colorado using the SnowModel snow evolution modeling system. SnowModel simulations were performed using a grid resolution of 100 m and 3-hourly time step over a 34-yr period (1984–2017). Atmospheric forcing was provided by phase 2 of the North American Land Data Assimilation System, and the simulations accounted for temporal changes in forest canopy from bark beetle and wildfire disturbances. Annual summary values of simulated snowpack properties [snow metrics; e.g., peak snow water equivalent (SWE), snowmelt rate and timing, and snow sublimation] were used to compute trends across the domain. Trends in simulated snow metrics varied depending on elevation, aspect, and land cover. Statistically significant trends did not occur evenly within the basin, and some areas were more sensitive than others. In addition, there were distinct trend differences between the different snow metrics. Upward trends in mean winter air temperature were 0.3°C decade −1 , and downward trends in winter precipitation were −52 mm decade −1 . Middle elevation zones, coincident with the greatest volumetric snow water storage, exhibited the greatest sensitivity to changes in peak SWE and snowmelt rate. Across the Rio Grande headwaters, snowmelt rates decreased by 20% decade −1 , peak SWE decreased by 14% decade −1 , and total snowmelt quantity decreased by 13% decade −1 . These snow trends are in general agreement with widespread snow declines that have been reported for this region. This study further quantifies these snow declines and provides trend information for additional snow variables across a greater spatial coverage at finer spatial resolution.

Colorado

Plausible combinations: An improved method to evaluate the covariate structure of Cormack-Jolly-Seber mark-recapture models

Mark-recapture models are extensively used in quantitative population ecology, providing estimates of population vital rates, such as survival, that are difficult to obtain using other methods. Vital rates are commonly modeled as functions of explanatory covariates, adding considerable flexibility to mark-recapture models, but also increasing the subjectivity and complexity of the modeling process. Consequently, model selection and the evaluation of covariate structure remain critical aspects of mark-recapture modeling. The difficulties involved in model selection are compounded in Cormack-Jolly- Seber models because they are composed of separate sub-models for survival and recapture probabilities, which are conceptualized independently even though their parameters are not statistically independent. The construction of models as combinations of sub-models, together with multiple potential covariates, can lead to a large model set. Although desirable, estimation of the parameters of all models may not be feasible. Strategies to search a model space and base inference on a subset of all models exist and enjoy widespread use. However, even though the methods used to search a model space can be expected to influence parameter estimation, the assessment of covariate importance, and therefore the ecological interpretation of the modeling results, the performance of these strategies has received limited investigation. We present a new strategy for searching the space of a candidate set of Cormack-Jolly-Seber models and explore its performance relative to existing strategies using computer simulation. The new strategy provides an improved assessment of the importance of covariates and covariate combinations used to model survival and recapture probabilities, while requiring only a modest increase in the number of models on which inference is based in comparison to existing techniques.

Open Journal Of Ecology