Geology ReportsSearch

SEARCH · Geology Reports

Results for “Journal of Agricultural, Biological, and Environmental Statistics”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A goodness-of-fit test for occupancy models with correlated within-season revisits

Occupancy modeling is important for exploring species distribution patterns and for conservation monitoring. Within this framework, explicit attention is given to species detection probabilities estimated from replicate surveys to sample units. A central assumption is that replicate surveys are independent Bernoulli trials, but this assumption becomes untenable when ecologists serially deploy remote cameras and acoustic recording devices over days and weeks to survey rare and elusive animals. Proposed solutions involve modifying the detection-level component of the model (e.g., first-order Markov covariate). Evaluating whether a model sufficiently accounts for correlation is imperative, but clear guidance for practitioners is lacking. Currently, an omnibus goodnessof- fit test using a chi-square discrepancy measure on unique detection histories is available for occupancy models (MacKenzie and Bailey, Journal of Agricultural, Biological, and Environmental Statistics, 9, 2004, 300; hereafter, MacKenzie– Bailey test). We propose a join count summary measure adapted from spatial statistics to directly assess correlation after fitting a model. We motivate our work with a dataset of multinight bat call recordings from a pilot study for the North American Bat Monitoring Program. We found in simulations that our join count test was more reliable than the MacKenzie–Bailey test for detecting inadequacy of a model that assumed independence, particularly when serial correlation was low to moderate. A model that included a Markov-structured detection-level covariate produced unbiased occupancy estimates except in the presence of strong serial correlation and a revisit design consisting only of temporal replicates. When applied to two common bat species, our approach illustrates that sophisticated models do not guarantee adequate fit to real data, underscoring the importance of model assessment. Our join count test provides a widely applicable goodness-of-fit test and specifically evaluates occupancy model lack of fit related to correlation among detections within a sample unit. Our diagnostic tool is available for practitioners that serially deploy survey equipment as a way to achieve cost savings.

Ecology and Evolution

Techniques to improve ecological interpretability of black box machine learning models

Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme

Controlling for varying effort in count surveys: An analysis of Christmas Bird Count data

The Christmas Bird Count (CBC) is a valuable source of information about midwinter populations of birds in the continental U.S. and Canada. Analysis of CBC data is complicated by substantial variation among sites and years in effort expended in counting; this feature of the CBC is common to many other wildlife surveys. Specification of a method for adjusting counts for effort is a matter of some controversy. Here, we present models for longitudinal count surveys with varying effort; these describe the effect of effort as proportional to exp(B effortp), where B and p are parameters. For any fixed p, our models are loglinear in the transformed explanatory variable (effort)p and other covariables. Hence, we fit a collection of loglinear models corresponding to a range of values of p and select the best effort adjustment from among these on the basis of fit statistics. We apply this procedure to data for six bird species in five regions, for the period 1959-1988.

Journal of Agricultural, Biological, and Environme

Analysis of the influence of spatial pattern in habitat selection studies

Design and analysis of wildlife habitat selection studies typically do not assess the effect of spatial pattern on the habitat selection process. Effects of landscape scale pattern on habitat selection cannot be accomplished without replicate study areas, because pattern is a single, albeit multifaceted, attribute of an area. For a single area, however, the influence of pattern-related characteristics, such as shape and edge shared with adjacent patches, can be estimated by using GLIM (McCullough and Neider 1983) procedures to model patch-specific frequency counts of animal use as a function of these parameters. This approach is evaluated and illustrated with simulated breeding-bird counts in a South Carolina study area for which a GIS land cover classification is available. A related technique for evaluating whether movement from patch to patch is selective is developed and illustrated for designs that involve collection of trajectory data from monitored individuals. These designs and analyses are feasible given current GIS and GPS technology. Statistical inferences from habitat selection studies should be interpreted within the context of a range of scales at which animals differentiate between patch attributes.

Journal of Agricultural, Biological, and Environme

Computationally efficient statistical differential equation modeling using homogenization

Statistical models using partial differential equations (PDEs) to describe dynamically evolving natural systems are appearing in the scientific literature with some regularity in recent years. Often such studies seek to characterize the dynamics of temporal or spatio-temporal phenomena such as invasive species, consumer-resource interactions, community evolution, and resource selection. Specifically, in the spatial setting, data are often available at varying spatial and temporal scales. Additionally, the necessary numerical integration of a PDE may be computationally infeasible over the spatial support of interest. We present an approach to impose computationally advantageous changes of support in statistical implementations of PDE models and demonstrate its utility through simulation using a form of PDE known as “ecological diffusion.” We also apply a statistical ecological diffusion model to a data set involving the spread of mountain pine beetle (Dendroctonus ponderosae) in Idaho, USA.

Journal of Agricultural, Biological, and Environme

Assessing first-order emulator inference for physical parameters in nonlinear mechanistic models

We present an approach for estimating physical parameters in nonlinear models that relies on an approximation to the mechanistic model itself for computational efficiency. The proposed methodology is validated and applied in two different modeling scenarios: (a) Simulation and (b) lower trophic level ocean ecosystem model. The approach we develop relies on the ability to predict right singular vectors (resulting from a decomposition of computer model experimental output) based on the computer model input and an experimental set of parameters. Critically, we model the right singular vectors in terms of the model parameters via a nonlinear statistical model. Specifically, we focus our attention on first-order models of these right singular vectors rather than the second-order (covariance) structure.

Journal of Agricultural, Biological, and Environme

Joint spatial modeling bridges the gap between disparate disease surveillance and population monitoring efforts informing conservation of at-risk bat species

White-Nose Syndrome (WNS) is a wildlife disease that has decimated hibernating bats since its introduction in North America in 2006. As the disease spreads westward, assessing the potentially differential impact of the disease on western bat species is an urgent conservation need. The statistical challenge is that the disease surveillance and species response monitoring data are not co-located, available at different spatial resolutions, non-Gaussian, and subject to observation error requiring a novel extension to spatially misaligned regression models for analysis. Previous work motivated by epidemiology applications has proposed two-step approaches that overcome the spatial misalignment while intentionally preventing the human health outcome from informing estimation of exposure. In our application, the impacted animals contribute to spreading the fungus that causes WNS, motivating development of a joint framework that exploits the known biological relationship. We introduce a Bayesian, joint spatial modeling framework that provides inferences about the impact of WNS on measures of relative bat activity and accounts for the uncertainty in estimation of WNS presence at non-surveyed locations. Our simulations demonstrate that the joint model produced more precise estimates of disease occurrence and unbiased estimates of the association between disease presence and the count response relative to competing two-step approaches. Our statistical framework provides a solution that leverages disparate monitoring activities and informs species conservation across large landscapes. Stan code and documentation are provided to facilitate access and adaptation for other wildlife disease applications.

Journal of Agricultural, Biological and Environmen

Bayesian approaches to proxy uncertainty quantification in paleoecology: A mathematical justification and practical integration

Paleoenvironmental data are essential for reconstructing environmental conditions in the distant past, and these reconstructions strongly depend on proxies and age–depth models. Proxies are indirect measurements that substitute for variables that cannot be directly measured, such as past precipitation. Conversely, an age–depth model is a tool that correlates the observed proxy with a specific moment in time. Bayesian age–depth modelling has proved to be a powerful method for estimating sediment ages and their associated uncertainties. However, there remains considerable potential for further integration into proxy analysis. In this paper, we explore a mathematical justification and a computational approach that integrates uncertainty at the age–depth level and propagates it to the proxy scale in the form of a posterior predictive distribution. This method mitigates potential biases and errors by removing the need to assign a single age to a given proxy measurement. It allows for quantifying the likelihood that proxy data values correspond to modelled ages, thus enabling the quantification of uncertainty in both the temporal and proxy value domains. The use of Bayesian statistics in proxy analysis represents a relatively recent advancement. We aim to mathematically justify incorporating the Markov chain Monte Carlo output from age–depth models into proxy analysis and to present a novel methodology for constructing environmental reconstructions using this approach.

Journal of Agricultural, Biological and Environmen

Assessing the fit of site-occupancy models

Few species are likely to be so evident that they will always be detected at a site when present. Recently a model has been developed that enables estimation of the proportion of area occupied, when the target species is not detected with certainty. Here we apply this modeling approach to data collected on terrestrial salamanders in the Plethodon glutinosus complex in the Great Smoky Mountains National Park, USA, and wish to address the question 'how accurately does the fitted model represent the data?' The goodness-of-fit of the model needs to be assessed in order to make accurate inferences. This article presents a method where a simple Pearson chi-square statistic is calculated and a parametric bootstrap procedure is used to determine whether the observed statistic is unusually large. We found evidence that the most global model considered provides a poor fit to the data, hence estimated an overdispersion factor to adjust model selection procedures and inflate standard errors. Two hypothetical datasets with known assumption violations are also analyzed, illustrating that the method may be used to guide researchers to making appropriate inferences. The results of a simulation study are presented to provide a broader view of the methods properties.

North Carolina, Tennessee

Extreme value-based methods for modeling elk yearly movements

Species range shifts and the spread of diseases are both likely to be driven by extreme movements, but are difficult to statistically model due to their rarity. We propose a statistical approach for characterizing movement kernels that incorporate landscape covariates as well as the potential for heavy-tailed distributions. We used a spliced distribution for distance travelled paired with a resource selection function to model movements biased toward preferred habitats. As an example, we used data from 704 annual elk movements around the Greater Yellowstone Ecosystem from 2001 to 2015. Yearly elk movements were both heavy-tailed and biased away from high elevations during the winter months. We then used a simulation to illustrate how these habitat effects may alter the rate of disease spread using our estimated movement kernel relative to a more traditional approach that does not include landscape covariates. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme

Limitations to mapping habitat use areas in changing landscapes using the Mahalanobis distance statistic

We tested the potential of a GIS mapping technique, using a resource selection model developed for black-tailed jackrabbits (Lepus californicus) and based on the Mahalanobis distance statistic, to track changes in shrubsteppe habitats in southwestern Idaho. If successful, the technique could be used to predict animal use areas, or those undergoing change, in different regions from the same selection function and variables without additional sampling. We determined the multivariate mean vector of 7 GIS variables that described habitats used by jackrabbits. We then ranked the similarity of all cells in the GIS coverage from their Mahalanobis distance to the mean habitat vector. The resulting map accurately depicted areas where we sighted jackrabbits on verification surveys. We then simulated an increase in shrublands (which are important habitats). Contrary to expectation, the new configurations were classified as lower similarity relative to the original mean habitat vector. Because the selection function is based on a unimodal mean, any deviation, even if biologically positive, creates larger Malanobis distances and lower similarity values. We recommend the Mahalanobis distance technique for mapping animal use areas when animals are distributed optimally, the landscape is well-sampled to determine the mean habitat vector, and distributions of the habitat variables does not change.

Journal of Agricultural, Biological, and Environme

Imputation approaches for animal movement modeling

The analysis of telemetry data is common in animal ecological studies. While the collection of telemetry data for individual animals has improved dramatically, the methods to properly account for inherent uncertainties (e.g., measurement error, dependence, barriers to movement) have lagged behind. Still, many new statistical approaches have been developed to infer unknown quantities affecting animal movement or predict movement based on telemetry data. Hierarchical statistical models are useful to account for some of the aforementioned uncertainties, as well as provide population-level inference, but they often come with an increased computational burden. For certain types of statistical models, it is straightforward to provide inference if the latent true animal trajectory is known, but challenging otherwise. In these cases, approaches related to multiple imputation have been employed to account for the uncertainty associated with our knowledge of the latent trajectory. Despite the increasing use of imputation approaches for modeling animal movement, the general sensitivity and accuracy of these methods have not been explored in detail. We provide an introduction to animal movement modeling and describe how imputation approaches may be helpful for certain types of models. We also assess the performance of imputation approaches in two simulation studies. Our simulation studies suggests that inference for model parameters directly related to the location of an individual may be more accurate than inference for parameters associated with higher-order processes such as velocity or acceleration. Finally, we apply these methods to analyze a telemetry data set involving northern fur seals ( Callorhinus ursinus ) in the Bering Sea. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme

The Bayesian group lasso for confounded spatial data

Generalized linear mixed models for spatial processes are widely used in applied statistics. In many applications of the spatial generalized linear mixed model (SGLMM), the goal is to obtain inference about regression coefficients while achieving optimal predictive ability. When implementing the SGLMM, multicollinearity among covariates and the spatial random effects can make computation challenging and influence inference. We present a Bayesian group lasso prior with a single tuning parameter that can be chosen to optimize predictive ability of the SGLMM and jointly regularize the regression coefficients and spatial random effect. We implement the group lasso SGLMM using efficient Markov chain Monte Carlo (MCMC) algorithms and demonstrate how multicollinearity among covariates and the spatial random effect can be monitored as a derived quantity. To test our method, we compared several parameterizations of the SGLMM using simulated data and two examples from plant ecology and disease ecology. In all examples, problematic levels multicollinearity occurred and influenced sampling efficiency and inference. We found that the group lasso prior resulted in roughly twice the effective sample size for MCMC samples of regression coefficients and can have higher and less variable predictive accuracy based on out-of-sample data when compared to the standard SGLMM.

Journal of Agricultural, Biological, and Environme

Estimator selection for closed-population capture: recapture

For valid statistical inference, it is important to select an appropriate statistical model. In the analysis of capture-recapture data under the closed-population models of Otis et al. (1978), information theoretic and hypothesis testing approaches to model selection are not practical, because some of the models have likelihoods with nonidenti- fiable parameters. A further problem is that, for some of the Otis et al. models, multiple estimators exist but there is no objective basis for deciding which estimator to use for a particular dataset. In CAPTURE, a computer program for estimating parameters un- der the closed models of Otis et al., a linear discriminant classifier is used to select an appropriate model. This classifier frequently selects the incorrect generating model in simulation studies, and it provides no guidance on which estimator to use once a model has been selected. In this study, we develop new classifiers for selecting the best esti- mator (as opposed to the generating model) and evaluate their performance. In addition, we investigate an estimator averaging approach to estimation that is a modification of the model averaging approach described by Buckland et al. (1997). We found that, in general, the overall performance of the new classifiers was unimpressive. In contrast, the estimator averaging approach we investigated performed well.

Journal of Agricultural, Biological, and Environme

Nonstationary demographic state-space models using unreplicated counts for species undergoing environmental stressors

A fundamental task in ecological statistics is to estimate abundance and growth rate distributions from wildlife monitoring data to inform conservation management. Modeling time series of wildlife populations presents a number of challenges from both statistical and ecological perspectives, including discreteness; lack of replication; nonstationarity; and observation, demographic, and other phenomenological processes. Nonstationary dynamics are often exhibited by populations undergoing environmental stressors. Models must account for these characteristics to produce reliable estimates of abundance and trends, yet estimation can be challenging with unreplicated data. We propose nonstationary demographic state-space models using unreplicated counts for populations undergoing environmental stressors. A reduced growth rate model matches the complexity of the unreplicated count data, and a fecundity bound on growth rate distributions allows the separation of processes affecting growth rates like environmental stressors from those affecting abundance external to growth rates like migration. NDSSMs allow for the embedding of nonstationary model components, and we explore the use of changepoints, volatility clustering, and migration processes. We apply the proposed nonstationary models in case studies of herons affected by predator/competitor reestablishment and three bat species affected by a fungal pathogen causing white-nose syndrome. Nonstationary models outperform stationary models and generalized linear mixed effects models according to model scoring and visual inspection of predictions, and provide estimates more consistent with published values. Incorporating migration improves model fit universally, even with approximate one-way immigration, most likely because populations are extirpated, recolonized, and increase multiple-fold over the upper bound set by species fecundity. In addition, estimates of the timing and severity of the environmental stressor differed for models with migration. Including nonstationary and demographic components in a fecundity-bounded growth rate model improves inference and benefits interpretability of hyperparameters. In turn, this adjusts uncertainties in predictions of abundance and growth rates over time, providing the ingredients needed for informed conservation analysis and for directing future monitoring of at-risk species.

Journal of Agricultural, Biological and Environmen

Ecological prediction with nonlinear multivariate time-frequency functional data models

Time-frequency analysis has become a fundamental component of many scientific inquiries. Due to improvements in technology, the amount of high-frequency signals that are collected for ecological and other scientific processes is increasing at a dramatic rate. In order to facilitate the use of these data in ecological prediction, we introduce a class of nonlinear multivariate time-frequency functional models that can identify important features of each signal as well as the interaction of signals corresponding to the response variable of interest. Our methodology is of independent interest and utilizes stochastic search variable selection to improve model selection and performs model averaging to enhance prediction. We illustrate the effectiveness of our approach through simulation and by application to predicting spawning success of shovelnose sturgeon in the Lower Missouri River.

Lower Missouri River

Insights into the latent multinomial model through mark-resight data on female grizzly bears with cubs-of-the-year

Mark-resight designs for estimation of population abundance are common and attractive to researchers. However, inference from such designs is very limited when faced with sparse data, either from a low number of marked animals, a low probability of detection, or both. In the Greater Yellowstone Ecosystem, yearly mark-resight data are collected for female grizzly bears with cubs-of-the-year (FCOY), and inference suffers from both limitations. To overcome difficulties due to sparseness, we assume homogeneity in sighting probabilities over 16 years of bi-annual aerial surveys. We model counts of marked and unmarked animals as multinomial random variables, using the capture frequencies of marked animals for inference about the latent multinomial frequencies for unmarked animals. We discuss undesirable behavior of the commonly used discrete uniform prior distribution on the population size parameter and provide OpenBUGS code for fitting such models. The application provides valuable insights into subtleties of implementing Bayesian inference for latent multinomial models. We tie the discussion to our application, though the insights are broadly useful for applications of the latent multinomial model.

Journal of Agricultural, Biological, and Environme

Hidden Markov model for dependent mark loss and survival estimation

Mark-recapture estimators assume no loss of marks to provide unbiased estimates of population parameters. We describe a hidden Markov model (HMM) framework that integrates a mark loss model with a Cormack–Jolly–Seber model for survival estimation. Mark loss can be estimated with single-marked animals as long as a sub-sample of animals has a permanent mark. Double-marking provides an estimate of mark loss assuming independence but dependence can be modeled with a permanently marked sub-sample. We use a log-linear approach to include covariates for mark loss and dependence which is more flexible than existing published methods for integrated models. The HMM approach is demonstrated with a dataset of black bears ( Ursus americanus ) with two ear tags and a subset of which were permanently marked with tattoos. The data were analyzed with and without the tattoo. Dropping the tattoos resulted in estimates of survival that were reduced by 0.005–0.035 due to tag loss dependence that could not be modeled. We also analyzed the data with and without the tattoo using a single tag. By not using.

Journal of Agricultural, Biological, and Environme