Geology ReportsSearch

Geology topics

Richard M. Mitchell

Publications and source records attributed to Richard M. Mitchell.

8 recordsLinked to original sources

Using models of local environmental conditions for biological assessment

A common approach for biological assessment is to compare current observations of biota at a site to predictions of the biota that would occur if the site were in reference condition. To estimate these reference expectations, predictive bioassessment models use only data collected at reference sites and estimate relationships between biota and environmental variables that are unaffected by human activities (i.e., immutable variables). However, in some areas where human activities are pervasive, few reference sites are available. Here, we introduce a new approach for bioassessment, in which we first estimate relationships between widely available measurements of both immutable and human-influenced landscape variables and local environmental conditions (conductivity, dissolved total P, total suspended solids, percentage of sand and fines in the substrate, and dissolved organic C), and then we estimate the relationship between these local environmental variables and total macroinvertebrate richness. We then calculate the values of the local environmental variables under reference conditions by adjusting predictors that are strongly influenced by human activities to levels that are consistent with those observed at reference sites. Reference values of local environmental variables are then combined with the model for taxon richness to calculate taxon richness expected under reference conditions. Predictions of total taxon richness in validation data using the new approach are similar in accuracy and precision as predictions calculated using a traditional approach that focuses only on reference site data. The new approach complements existing bioassessment methods by highlighting types of sites in which estimates of reference expectations are uncertain and by improving our understanding of how instream stressors affect expected total taxon richness.

conterminous United States

A web-based tool for assessing the condition of benthic diatom assemblages in streams and rivers of the conterminous United States

Benthic diatom assemblages are known to be indicative of water quality but have yet to be widely adopted in biological assessments in the United States due to several limitations. Our goal was to address some of these limitations by developing regional multi-metric indices (MMIs) that are robust to inter-laboratory taxonomic inconsistency, adjusted for natural covariates, and sensitive to a wide range of anthropogenic stressors. We aggregated bioassessment data from two national-scale federal programs and used a data-driven analysis in which all-possible combinations of 2–7 metrics were compared for three measures of performance. After ranking the best-performing MMIs, we selected the final MMIs by evaluating stress-response relations in independent regional datasets of diatom samples paired with measures of several water-quality stressors, including herbicides and streamflow flashiness. Each regional MMI performed well at calibration sites and represented diverse aspects of the structure and function of diatom communities. Most metrics included in the best MMIs were modeled to account for natural variation including climate, topography, soil characteristics, lithology, and groundwater influence on streamflow. MMI performance improved with higher numbers of component metrics, but this effect diminished beyond six metrics. Component metrics of MMIs were associated with a broad suite of measured stressors in every region, including salinity, nutrients, herbicides, and streamflow flashiness. We provide a web-based software application that allows users in the conterminous United States to apply our MMIs to their own datasets and compare MMI scores from their sites to a broader regional context.

Ecological Indicators

Techniques to improve ecological interpretability of black box machine learning models

Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme

Reduction of taxonomic bias in diatom species data

Inconsistency in taxonomic identification and analyst bias impede the effective use of diatom data in regional and national stream and lake surveys. In this study, we evaluated the effect of existing protocols and a revised protocol on the precision of diatom species counts. The revised protocol adjusts four elements of sample preparation, taxon identification and enumeration, and quality control (QC). We used six independent datasets to assess the effect of the adjustments on analytical outcomes. The first dataset was produced by three laboratories with a total of five analysts following established protocols (Charles et al. 2002), or their slight variations. The remaining datasets were produced by 1-3 laboratories with a total of 2-3 analysts following a revised protocol. The revised protocol included the following modifications: 1) development of coordinated pre-count voucher floras based on morphological operational taxonomic units (mOTUs), 2) random assignment of samples to analysts, 3) post-count identification and documentation of taxa (as opposed to an approach in which analysts assign names while they enumerate), and 4) increased use of QC samples. The revised protocol reduced taxonomic bias, as measured by reduction in analyst signal, and improved similarity among QC samples. Reduced taxonomic bias improves the performance of biological assessments, facilitates transparency across studies, and refines estimates of diatom species distributions.

Limnology and Oceanography: Methods

A random forest approach for bounded outcome variables

Random forests have become an established tool for classication and regres- sion, in particular in high-dimensional settings and in the presence of non-additive predictor-response relationships. For bounded outcome variables restricted to the unit interval, however, classical modeling approaches based on mean squared error loss may severely suer as they do not account for heteroscedasticity in the data. To address this issue, we propose a random forest approach for relating a beta dis- tributed outcome to a set of explanatory variables. Our approach explicitly makes use of the likelihood function of the beta distribution for the selection of splits dur- ing the tree-building procedure. In each iteration of the tree-building algorithm it chooses one explanatory variable in combination with a split point that maximizes the log-likelihood function of the beta distribution with the parameter estimates de- rived from the nodes of the currently built tree. Results of several simulation studies and an application using data from the U.S.A. National Lakes Assessment Survey demonstrate the properties and usefulness of the method, in particular when com- pared to random forest approaches based on mean squared error loss and parametric regression models.

Journal of Computational and Graphical Statistics

Assessing water-quality changes in U.S. rivers at multiple geographic scales using results from probabilistic and targeted monitoring

Two commonly used approaches for water quality monitoring are probabilistic and targeted. In a probabilistic approach like the US Environmental Protection Agency’s National Rivers and Streams Assessment, monitoring sites are selected using a statistically representative approach. In a targeted approach like that used by many monitoring organizations, monitoring sites are chosen individually to answer specific questions. One important goal of both approaches is documenting long-term changes in water quality. Here, we compare chloride change results in US rivers and streams between the early 2000s and early 2010s from both approaches. The probabilistic approach provided an unbiased representation of change in all US rivers and streams, but was designed to measure low-streamflow conditions within a spring/summer index period during periodic survey years. The targeted approach was focused on larger, more developed watersheds but samples were collected frequently throughout the assessment period in different seasons and streamflows. The probabilistic results showed a small decrease in chloride concentrations in rivers and streams with the lowest concentrations, but no consistent increase or decrease in the remainder. The increased granularity of the targeted results showed that there was, in fact, a mix of changes occurring, with increases at 132 sites, decreases at 112 sites, and relatively stable conditions at 55 sites. The combined results suggest that chloride is not responding to a widespread, common driver across the USA and that management of chloride would be most effective when targeted regionally or locally.

Environmental Monitoring and Assessment

Taxonomic harmonization may reveal a stronger association between diatom assemblages and total phosphorus in large datasets

Diatom data have been collected in large-scale biological assessments in the United States, such as the U.S. Environmental Protection Agency’s National Rivers and Streams Assessment (NRSA). However, the effectiveness of diatoms as indicators may suffer if inconsistent taxon identifications across different analysts obscure the relationships between assemblage composition and environmental variables. To reduce these inconsistencies, we harmonized the 2008–2009 NRSA data from nine analysts by updating names to current synonyms and by statistically identifying taxa with high analyst signal (taxa with more variation in relative abundance explained by the analyst factor, relative to environmental variables). We then screened a subset of samples with QA/QC data and combined taxa with mismatching identifications by the primary and secondary analysts. When these combined “slash groups” did not reduce analyst signal, we elevated taxa to the genus level or omitted taxa in difficult species complexes. We examined the variation explained by analyst in the original and revised datasets. Further, we examined how revising the datasets to reduce analyst signal can reduce inconsistency, thereby uncovering the variation in assemblage composition explained by total phosphorus (TP), an environmental variable of high priority for water managers. To produce a revised dataset with the greatest taxonomic consistency, we ultimately made 124 slash groups, omitted 7 taxa in the small naviculoid (e.g., Sellaphora atomoides ) species complex, and elevated Nitzschia , Diploneis , and Tryblionella taxa to the genus level. Relative to the original dataset, the revised dataset had more overlap among samples grouped by analyst in ordination space, less variation explained by the analyst factor, and more than double the variation in assemblage composition explained by TP. Elevating all taxa to the genus level did not eliminate analyst signal completely, and analyst remained the most important predictor for the genera Sellaphora , Mayamaea , and Psammodictyon , indicating that these taxa present the greatest obstacle to consistent identification in this dataset. Although our process did not completely remove analyst signal, this work provides a method to minimize analyst signal and improve detection of diatom association with TP in large datasets involving multiple analysts. Examination of variation in assemblage data explained by analyst and taxonomic harmonization may be necessary steps for improving data quality and the utility of diatoms as indicators of environmental variables.

Ecological Indicators

Land use patterns, ecoregion, and microcystin relationships in U.S. lakes and reservoirs: a preliminary evaluation

A statistically significant association was found between the concentration of total microcystin, a common class of cyanotoxins, in surface waters of lakes and reservoirs in the continental U.S. with watershed land use using data from 1156 water bodies sampled between May and October 2007 as part of the USEPA National Lakes Assessment. Nearly two thirds (65.8%) of the samples with microcystin concentrations ≥1.0 μg/L (n = 126) were limited to three nutrient and water quality-based ecoregions (Corn Belt and Northern Great Plains, Mostly Glaciated Dairy Region, South Central Cultivated Great Plains) in watersheds with strong agricultural influence. canonical correlation analysis (CCA) indicated that both microcystin concentrations and cyanobacteria abundance were positively correlated with total nitrogen, dissolved organic carbon, and temperature; correlations with total phosphorus and water clarity were not as strong. This study supports a number of regional lake studies that suggest that land use practices are related to cyanobacteria abundance, and extends the potential impacts of agricultural land use in watersheds to include the production of cyanotoxins in lakes.

Harmful Algae