Geology ReportsSearch

SEARCH · Geology Reports

Results for “Machine Learning with Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Seismology with dark data: Image-based processing of analog records using machine learning for the rangely earthquake control experiment

Before the digital era, seismograms were recorded in analog form and read manually by analysts. The digital era represents only about 25% of the total time span of instrumental seismology. Analog data provide important constraints on earthquake processes over the long term, and in some cases are the only data available. The media on which analog data are recorded degrades with time and there is an urgent need for cost‐effective approaches to preserve the information they contain. In this study, we work directly with images by constructing a set of image‐based methods for earthquake processing, rather than pursue the usual approach of converting analog data to vector time series. We demonstrate this approach on one month of continuous Develocorder films from the Rangely earthquake control experiment run by the U.S. Geological Survey (USGS). We scan the films into images and compress these into low‐dimensional feature vectors as input to a classifier that separates earthquakes from noise in a defined feature space. We feed the detected event images into a short‐term average/long‐term average (STA/LTA) picker, a grid‐search associator, and a 2D image correlator to measure both absolute arrival times and relative arrival‐time differences between events. We use these measurements to locate the earthquakes using hypoDD. In the month that we studied, we identified 40 events clustered near the injection wells. In the original study, Raleigh et al. (1976) identified only 32 events during the same period. Scanning without vectorizing analog seismograms represents an attractive approach to archiving these perishable data. We demonstrated that it is possible to carry out precision seismology directly on such images. Our approach has the potential for wide application to analog seismograms.

Seismological Research Letters

Human-in-the-Loop segmentation of earth surface imagery

Segmentation, or the classification of pixels (grid cells) in imagery, is ubiquitously applied in the natural sciences. Manual methods are often prohibitively time-consuming, especially those images consisting of small objects and/or significant spatial heterogeneity of colors or textures. Labeling complicated regions of transition that in Earth surface imagery are represented by collections of mixed-pixels, -textures, and -spectral signatures, can be especially error-prone because it is difficult to reliably unmix, identify and delineate consistently. However, the success of supervised machine learning (ML) approaches is entirely dependent on good label data. We describe a fast, semi-automated, method for interactive segmentation of N-dimensional (x, y, N) images into two-dimensional (x, y) label images. It uses human-in-the-loop ML to achieve consensus between the labeler and a model in an iterative workflow. The technique is reproducible; the sequence of decisions made by human labeler and ML algorithms can be encoded to file, so the entire process can be played back and new outputs generated with alternative decisions and/or algorithms. We illustrate the scientific potential of segmentation of imagery of diverse settings and image types using six case studies from river, estuarine, and open coast environments. These photographic and non-photographic imagery consist of 1- and 3-bands on regular and irregular grids ranging from centimeters to tens of meters. We demonstrate high levels of agreement in label images generated by several labelers on the same imagery, and make suggestions to achieve consensus and measure uncertainty, ideal for widespread application in training supervised ML for image segmentation.

Earth and Space Science

An integrated sensor network and data driven approach to satellite remote sensing of dissolved organic matter

Traditional remote sensing retrieval models for water quality have historically relied on limited, localized data sets due to the prohibitive costs of extensive field campaigns and logistical challenges of collecting match-up data with satellite overpasses. As a result, these models often lack generalizability across seasons, tides, and sites. Furthermore, small field data sets limit the utility of modern machine learning techniques to advance remote sensing retrieval models. In situ optical sensors deployed in a sensor network to continuously monitor larger water bodies can drastically increase the number of measurements, providing the opportunity to develop new approaches for building robust remote sensing retrieval models by leveraging both remote sensing data and in situ networks as an integrated monitoring system. This study leverages a large “ground-to-space” sensor network that combines an in situ optical sensor network with satellite-based remote sensing to overcome these limitations. Utilizing a large-scale data set from the U.S. Geological Survey's Sacramento—San Joaquin River Delta monitoring network, of dissolved organic matter fluorescence measurements, and remote sensing data from the European Space Agency's Sentinel-2A and -2B satellites, this study implemented a data driven approach for dissolved organic matter models. The data set, consisting of 982 samples collected between 2018 and 2021 was used to train and validate a random forest model ( R 2 = 0.76, RMSE = 6.1 Quinine Sulfate Equivalents), with demonstrated applicability across diverse site conditions, tidal stages, and seasons. This work provides a scalable solution to address critical challenges in water quality monitoring and offers a replicable framework for global water quality management.

Calfornia

GeoNat v1.0: A dataset for natural feature mapping with artificial intelligence and supervised learning

Machine learning allows “the machine” to deduce the complex and sometimes unrecognized rules governing spatial systems, particularly topographic mapping, by exposing it to the end product. Often, the obstacle to this approach is the acquisition of many good and labeled training examples of the desired result. Such is the case with most types of natural features. To address such limitations, this research introduces GeoNat v1.0, a natural feature dataset, used to support artificial intelligence‐based mapping and automated detection of natural features under a supervised learning paradigm. The dataset was created by randomly selecting points from the U.S. Geological Survey’s Geographic Names Information System and includes approximately 200 examples each of 10 classes of natural features. Resulting data were tested in an object‐detection problem using a region‐based convolutional neural network. The object‐detection tests resulted in a 62% mean average precision as baseline results. Major challenges in developing training data in the geospatial domain, such as scale and geographical representativeness, are addressed in this article. We hope that the resulting dataset will be useful for a variety of applications and shed light on training data collection and labeling in the geospatial artificial intelligence domain.

Transactions in GIS

Preliminary machine learning models of manganese and 1,4-dioxane in groundwater on Long Island, New York

Manganese and 1,4-dioxane in groundwater underlying Long Island, New York, were modeled with machine learning methods to demonstrate the use of these methods for mapping contaminants in groundwater in the Long Island aquifer system. XGBoost, a gradient boosted, ensemble tree method, was applied to data from 910 wells for manganese and 553 wells for 1,4-dioxane. Explanatory variables included soil properties, groundwater flow, land use, and other features that describe the hydrogeology and geochemistry of the aquifer system. Four models were developed to predict the probability of manganese concentrations greater than a detection level of 10 micrograms per liter (μg/L) and greater than three threshold concentrations (50, 150, and 300 μg/L) relevant to drinking-water quality. One model was developed to predict the probability of 1,4-dioxane concentrations greater than a detection level of 0.07 μg/L. The 1,4-dioxane model was limited geographically to Suffolk County because of data availability. Predictions were made for two layers in the upper glacial aquifer and three layers in the Magothy aquifer, which are the upper two of the three major aquifers of the Long Island aquifer system. The objective of the study described in this report was to demonstrate the application of the methods rather than to develop precise estimates of manganese or 1,4-dioxane concentrations at any given location. The predictive models developed in the study are considered preliminary in the sense that they are an initial effort at developing these kinds of models specifically for Long Island. The models could be improved by the inclusion of additional data, by the use of methods to improve the modeling of infrequent high concentrations of manganese and 1,4-dioxane (above threshold concentrations), and by including more explanatory variables that specifically describe conditions and contaminant sources on Long Island. Nonetheless, the distribution of model predictions and the influence of explanatory variables in the models were consistent with the expected relations between contaminant concentrations and groundwater-flow-system characteristics and the distribution of manmade sources. Mapped predictions indicated that manganese detections were more probable in the upper glacial aquifer and along the southern shore of Long Island, consistent with the distribution of anoxic conditions in groundwater in the Long Island aquifer system. Manganese was infrequently predicted at concentrations greater than thresholds of concern for drinking-water quality in any of the aquifer layers. Detections of 1,4-dioxane were predicted in the western, more highly developed parts of Suffolk County, in the upper glacial aquifer and the top and middle layers of the Magothy aquifer, and in northwestern Suffolk County in the bottom layer of the Magothy aquifer. Although preliminary in nature and based on limited data, these mapped predictions can be used to generally identify areas where manganese and 1,4-dioxane may be present at concentrations of concern to prioritize areas for future monitoring and to guide future modeling and mapping efforts.

New York

Machine learning can assign geologic basin to produced water samples using major ion geochemistry

Understanding the geochemistry of waters produced during petroleum extraction is essential to informing the best treatment and reuse options, which can potentially be optimized for a given geologic basin. Here, we used the US Geological Survey’s National Produced Waters Geochemical Database (PWGD) to determine if major ion chemistry could be used to classify accurately a produced water sample to a given geologic basin based on similarities to a given training dataset. Two datasets were derived from the PWGD: one with seven features but more samples (PWGD7), and another with nine features but fewer samples (PWGD9). The seven-feature dataset, prior to randomly generating a training and testing (i.e., validation) dataset, had 58,541 samples, 20 basins, and was classified based on total dissolved solids (TDS), bicarbonate (HCO 3 ), Ca, Na, Cl, Mg, and sulfate (SO 4 ). The nine-feature dataset, prior to randomly splitting into a training and testing (i.e., validation) dataset, contained 33,271 samples, 19 basins, and was classified based on TDS, HCO 3 , Ca, Na, Cl, Mg, SO 4 , pH, and specific gravity. Three supervised machine learning algorithms—Random Forest, k-Nearest Neighbors, and Naïve Bayes—were used to develop multi-class classification models to predict a basin of origin for produced waters using major ion chemistry. After training, the models were tested on three different datasets: Validation7, Validation9, and one based on data absent from the PWGD. Prediction accuracies across the models ranged from 23.5 to 73.5% when tested on the two PWGD-based datasets. A model using the Random Forest algorithm predicted most accurately compared to all other models tested. The models generally predicted basin of origin more accurately on the PWGD7-based dataset than on the PWGD9-based dataset. An additional dataset, which contained data not in the PWGD, was used to test the most accurate model; results suggest that some basins may lack geochemical diversity or may not be well described, while others may be geochemically diverse or are well described. A compelling result of this work is that a produced water basin of origin can be determined using major ions alone and, therefore, deep basinal fluid compositions may not be as variable within a given basin as previously thought. Applications include predicting the geochemistry of produced fluid prior to drilling at different intervals and assigning historical produced water data to a producing basin.

Natural Resources Research

A multimodal data fusion and deep learning framework for large-scale wildfire surface fuel mapping

Accurate estimation of fuels is essential for wildland fire simulations as well as decision-making related to land management. Numerous research efforts have leveraged remote sensing and machine learning for classifying land cover and mapping forest vegetation species. In most cases that focused on surface fuel mapping, the spatial scale of interest was smaller than a few hundred square kilometers; thus, many small-scale site-specific models had to be created to cover the landscape at the national scale. The present work aims to develop a large-scale surface fuel identification model using a custom deep learning framework that can ingest multimodal data. Specifically, we use deep learning to extract information from multispectral signatures, high-resolution imagery, and biophysical climate and terrain data in a way that facilitates their end-to-end training on labeled data. A multi-layer neural network is used with spectral and biophysical data, and a convolutional neural network backbone is used to extract the visual features from high-resolution imagery. A Monte Carlo dropout mechanism was also devised to create a stochastic ensemble of models that can capture classification uncertainties while boosting the prediction performance. To train the system as a proof-of-concept, fuel pseudo-labels were created by a random geospatial sampling of existing fuel maps across California. Application results on independent test sets showed promising fuel identification performance with an overall accuracy ranging from 55% to 75%, depending on the level of granularity of the included fuel types. As expected, including the rare—and possibly less consequential—fuel types reduced the accuracy. On the other hand, the addition of high-resolution imagery improved classification performance at all levels.

California

Interrogating process deficiencies in large-scale hydrologic models with interpretable machine learning

Large-scale hydrologic models are increasingly being developed for operational use in the forecasting and planning of water resources. However, the predictive strength of such models depends on how well they resolve various functions of catchment hydrology, which are influenced by gradients in climate, topography, soils, and land use. Most assessments of hydrologic model uncertainty have been limited to traditional statistical methods. Here, we present a proof-of-concept approach that uses interpretable machine learning techniques to provide post hoc assessment of model sensitivity and process deficiency in hydrologic models. We train a random forest model to predict the Kling–Gupta efficiency (KGE) of National Water Model (NWM) and National Hydrologic Model (NHM) streamflow predictions for 4383 stream gauges in the conterminous United States. Thereafter, we explain the local and global controls that 48 catchment attributes exert on KGE prediction using interpretable Shapley values. Overall, we find that soil water content is the most impactful feature controlling successful model performance, suggesting that soil water storage is difficult for hydrologic models to resolve, particularly for arid locations. We identify nonlinear thresholds beyond which predictive performance decreases for NWM and NHM. For example, soil water content less than 210 mm, precipitation less than 900 mm yr −1 , road density greater than 5 km km −2 , and lake area percent greater than 10 % contributed to lower KGE values. These results suggest that improvements in how these influential processes are represented could result in the largest increases in NWM and NHM predictive performance. This study demonstrates the utility of interrogating process-based models using data-driven techniques, which has broad applicability and potential for improving the next generation of large-scale hydrologic models.

conterminous United States

Beyond overlap: Considering habitat preference and fitness outcomes in the umbrella species concept

Umbrella species and other surrogate species approaches to conservation provide an appealing framework to extend the reach of conservation efforts beyond single species. For the umbrella species concept to be effective, populations of multiple species of concern must persist in areas protected on behalf of the umbrella species. Most assessments of the concept, however, focus exclusively on geographic overlap among umbrella and background species, and not measures that affect population persistence (e.g. habitat quality or fitness). We quantified the congruence between the habitat preferences and nesting success of a high-profile umbrella species (greater sage-grouse, Centrocercus urophasianus , hereafter ‘sage-grouse’), and three sympatric species of declining songbirds (Brewer's sparrow Spizella breweri , sage thrasher Oreoscoptes montanus and vesper sparrow Pooecetes gramineus ) in central Wyoming, USA during 2012–2013. We used machine-learning methods to create data-driven predictions of sage-grouse nest-site selection and nest survival probabilities by modeling field-collected sage-grouse data relative to habitat attributes. We then used field-collected songbird data to assess whether high-quality sites for songbirds aligned with those of sage-grouse. Nest sites selected by songbirds did not coincide with sage-grouse nesting preferences, with the exception that Brewer's sparrows preferred similar nest sites to sage-grouse in 2012. Moreover, the areas that produced higher rates of songbird nest survival were unrelated to those for sage-grouse. Our findings suggest that management actions at local scales that prioritize sage-grouse nesting habitat will not necessarily enhance the reproductive success of sagebrush-associated songbirds. Measures implemented to conserve sage-grouse and other purported umbrella species at broad spatial scales likely overlap the distribution of many species, however, broad-scale overlap may not translate to fine-scale conservation benefit beyond the umbrella species itself. The maintenance of microhabitat heterogeneity important for a diversity of species of concern will be critical for a more holistic application of the umbrella species concept.

Animal Conservation

Lava lake thermal pattern classification using self organizing maps and relationships to eruption processes at Kilauea Volcano, Hawaii

Kīlauea Volcano’s active summit lava lake poses hazards to downwind residents and over 1.6 million Hawai‘i Volcanoes National Park visitors each year. The lava lake surface is dynamic; crustal plates separated by incandescent cracks move across the lake as magma circulates below. We hypothesize that these dynamic thermal patterns are related to changes in other volcanic processes, such that sequences of thermal images may provide information about eruption parameters that are sometimes difficult to measure. The ability to learn about current gas emissions and seismic activity from a remote thermal time-lapse camera would be beneficial when conditions are too hazardous for field measurements. We apply a machine learning algorithm called self-organizing maps (SOM) to thermal infrared time-lapse images of the lava lake collected hourly over 23 April – 21 October 2013 (n=4354). The SOM algorithm can take thousands of seemingly different images, each representing the spatial distribution of relative temperature across the lava lake surface, and group them into clusters based on their similarities. We then relate the resulting clusters to sulfur dioxide emissions and seismic tremor to characterize ties between the SOM classification and different emplacement conditions. The SOM classification results are highly sensitive to the normalization method applied to the input images. The standard pixel-by-pixel normalization method yields a cluster of images defined by the highest observed SO2 emission levels, elevated surface temperatures, and a high proportion of cracks between crustal plates. When lava lake surface patterns are isolated by minimizing the effect of temperature variation between images, relationships with seismic tremor activity emerge, revealing an “intense spatter” cluster, characterized by unstable, broken-up crustal plate patterns on the lava lake surface. This proof of concept study provides a basis for extending the SOM classification method to hazard forecasting and real-time volcanic monitoring applications, as well as comparative studies at other lava lakes.

Hawaii

Development and application of a coastal change likelihood assessment for the northeast region, Maine to Virginia

Coastal resources are increasingly affected by erosion, extreme weather events, sea level rise, tidal flooding, and other potential hazards related to climate change. These hazards have varying effects on coastal landscapes because of the compounding of geologic, oceanographic, ecologic, and socioeconomic factors that exist at a given location. An assessment framework is introduced in this report that synthesizes existing datasets that cover the variability of the landscape, and hazards that may act on the landscape, to evaluate the likelihood of coastal change along the U.S. coastline on a decadal scale. The pilot study that aided in the development of the framework was run in the northeastern United States (from Maine to Virginia) and consists of datasets derived from a variety of Federal, State, and local sources. First, a decision-tree-based dataset was built that describes the resistance or integrity of the coastal landscape (called the fabric dataset for the purposes of this report) and includes land cover, elevation, slope, long-term (more than 50 years) shoreline change, dune height, and marsh stability data. A second database was generated from coastal hazards, which are divided into event hazards (for example, flooding, wave power, and probability of storm overwash) and persistent or perpetual hazards (for example, relative sea level rise rate, short-term [about 30-year] shoreline erosion rate, and storm recurrence interval). The fabric dataset was then merged with the coastal hazards databases, and a model training dataset made up of hundreds of polygons was generated from these combined data to support machine learning. The pilot study resulted in location-specific, 10-meter-resolution data classified into five raster datasets that include intrinsic characteristics of the coast used to determine the resistance of the landscape to change, the persistent and event hazards that act on the coast, the machine learning output (coastal change likelihood) based on the cumulative effects of the fabric and hazards datasets, and an estimate of the hazard type (event or persistent) that is the most likely to influence coastal change. Final outcomes are intended to be used as a first-order planning tool to determine which areas of the coast are more likely to change in response to future potential coastal hazards and to examine elements and drivers that make change in a location more likely.

Connecticut, Delaware, Maine, Maryland, Massachuse

Machine learning predictions of nitrate in groundwater used for drinking supply in the conterminous United States

Groundwater is an important source of drinking water supplies in the conterminous United State (CONUS), and presence of high nitrate concentrations may limit usability of groundwater in some areas because of the potential negative health effects. Prediction of locations of high nitrate groundwater is needed to focus mitigation and relief efforts. A three-dimensional extreme gradient boosting (XGB) machine learning model was developed to predict the distribution of nitrate. Nitrate was predicted at a 1 km resolution for two drinking water zones, each of variable depth, one for domestic supply and one for public supply. The model used measured nitrate concentrations from 12,082 wells and included predictor variables representing well characteristics, hydrologic conditions, soil type, geology, land use, climate, and nitrogen inputs. Predictor variables derived from empirical or numerical process-based models were also included to integrate information on controlling processes and conditions. The model provided accurate estimates at national and regional scales: the training (R 2 of 0.83) and hold-out (R 2 of 0.49) data fits compared favorably to previous studies. Predicted nitrate concentrations were less than 1 mg/L across most of the CONUS. Nationally, well depth, soil and climate characteristics, and the absence of developed land use were among the most influential explanatory factors. Only 1% of the area in either water supply zone had predicted nitrate concentrations greater than 10 mg/L; however, about 1.4 M people depend on groundwater for their drinking supplies in those areas. Predicted high concentrations of nitrate were most prevalent in the central CONUS. In areas of predicted high nitrate concentration, applied manure, farm fertilizer , and agricultural land use were influential predictor variables. This work represents the first application of XGB to a three-dimensional national-scale groundwater quality model and provides a significant milestone in the efforts to document nitrate in groundwater across the CONUS.

Science of the Total Environment

Hyperspectral imaging of river bathymetry using an ensemble of regression trees

Remote sensing has emerged as an effective tool for characterizing river systems, and machine learning (ML) techniques could make this approach even more powerful. To explore this possibility, we developed an ML-based workflow for hyperspectral imaging of river bathymetry using an ensemble of regression trees (HIRBERT). This approach involves using paired observations of depth and reflectance to select wavelength bands as predictors and then train a depth retrieval model; applying the model to the image yields a spatially continuous bathymetric map. We used data from five rivers with diverse morphologies and optical characteristics to assess whether HIRBERT can (1) provide more accurate depth estimates than a band ratio-based algorithm and (2) extend the range of depths detectable via remote sensing. Relative to single band combinations identified via optimal band ratio analysis (OBRA), regression tree ensembles improved depth retrieval performance, with observed versus predicted (OP) regression R 2 values increasing for all five sites. Similarly, HIRBERT provided more reliable depth estimates than OBRA over the full range of depths present along each river. These results suggest that by incorporating additional spectral information from multiple wavelength bands, ML could enhance bathymetric mapping across a range of river environments. In addition, we show how graphical tools can facilitate interpretation of ML-based depth retrieval models and yield insight regarding relationships between depth and reflectance. The HIRBERT workflow is packaged in free, standalone software developed to support applications in river research and management. Although ML can enhance remote sensing of river bathymetry, the limitations of this approach must also be acknowledged: Field measurements of water depth are required to train a depth retrieval model and the resulting model should only be applied to the image from which the training data were derived. The inherently image-specific nature of this approach implies that developing generalized regression tree ensembles that could be applied at larger scales would require additional research.

California, Idaho, Nebraska, Oregon, Wyoming

A generalized deep learning model to detect and classify volcano seismicity

Volcano seismicity is often detected and classified based on its spectral properties. However, the wide variety of volcano seismic signals and increasing amounts of data make accurate, consistent, and efficient detection and classification challenging. Machine learning (ML) has proven very effective at detecting and classifying tectonic seismicity, particularly using Convolutional Neural Networks (CNNs) and leveraging labeled datasets from regional seismic networks. Progress has been made applying ML to volcano seismicity, but efforts have typically been focused on a single volcano and are often hampered by the limited availability of training data. We build on the method of Tan et al. [2024] ( 10.1029/2024JB029194 ) to generalize a spectrogram-based CNN termed the VOlcano Infrasound and Seismic Spectrogram Neural Network ( VOISS-Net ) to detect and classify volcano seismicity at any volcano. We use a diverse training dataset of over 270,000 spectrograms from multiple volcanoes: Pavlof, Semisopochnoi, Tanaga, Takawangha, and Redoubt volcanoes\replaced (Alaska, USA); Mt. Etna (Italy); and Kīlauea, Hawai`i (USA). These volcanoes present a wide range of volcano seismic signals, source-receiver distances, and eruption styles. Our generalized VOISS-Net model achieves an accuracy of 87 % on the test set. We apply this model to continuous data from several volcanoes and eruptions included within and outside our training set, and find that multiple types of tremor, explosions, earthquakes, long-period events, and noise are successfully detected and classified. The model occasionally confuses transient signals such as earthquakes and explosions and misclassifies seismicity not included in the training dataset (e.g. teleseismic earthquakes). We envision the generalized VOISS-Net model to be applicable in both research and operational volcano monitoring settings.

Volcanica

Automated snow cover detection on mountain glaciers usingspaceborne imagery and machine learning

Tracking the extent of seasonal snow on glaciers over time is critical for assessing glacier vulnerability and the response of glacierized watersheds to climate change. Existing snow cover products do not reliably distinguish seasonal snow from glacier ice and firn, preventing their use for glacier snow cover detection. Despite previous efforts to classify glacier surface facies using machine learning on local scales, currently there is no published comparison of machine learning models for classifying glacier snow cover across different satellite image products. We present an automated snow detection workflow for mountain glaciers using supervised machine-learning-based image classifiers and Landsat 8 and 9, Sentinel-2, and PlanetScope satellite imagery. We develop the image classifiers by testing numerous machine learning algorithms with training and validation data from the U.S. Geological Survey Benchmark Glacier Project glaciers. The workflow produces daily to twice monthly time series of several glacier mass balance and snowmelt indicators (snow-covered area, accumulation area ratio, and seasonal snow line) from 2013 to present. Workflow performance is assessed by comparing automatically classified images and snow lines to manual interpretations at each glacier site. The image classifiers exhibit overall accuracies of 92%–98%, K scores of 84%–96%, and F scores of 93%–98% for all image products. The median difference between automatically and manually delineated median snow line altitudes is 31m (IQR of 73to0m)across all image products. The Sentinel-2 classifier (support vector machine) produces the most accurate glacier mass balance and snowmelt indicators and distinguishes snow from ice and f irn the most reliably. Although they are less accurate, the Landsat- and PlanetScope-derived estimates greatly enhance the temporal coverage of observations. The transient accumulation area ratio produces the least noisy time series, making it the most reliable indicator for characterizing seasonal snow trends. The temporally detailed accumulation area ratio time series reveal that the timing of minimum snow cover conditions varies by up to a month between Arctic (63°N) and midlatitude (48°N) sites, underscoring the potential for bias when estimating glacier minimum snow cover conditions from a single late-summer image. Widespread application of our automated snow detection workflow has the potential to improve regional assessments of glacier mass balance, land ice representations within Earth system models, water resources, and the impacts of climate change on snow cover across broad spatial scales.

The Cryosphere

A transferable approach for quantifying benthic fish sizes and densities in annotated underwater images

1. Benthic fishes are a common target of scientific monitoring but are difficult to quantify because of their close association to bottom habitats that are hard to access. Advances in image-acquisition technologies, machine vision, and deep learning have made capturing and quantifying fishes with cameras increasingly feasible. We present a method and open-source software called ‘FishScale’ to estimate benthic fish lengths, numeric abundance, and biomass density in underwater environments assessed with down-looking monocular images. 2. ‘FishScale’ estimates fish abundances and size frequencies from near-nadir monocular images where fish have already been semantically segmented. The software accounts for lens distortion, underwater magnification effects, and fish body curvature to automatically estimate fish lengths and the areas of images where they were captured. Numeric and biomass density are estimated through a deterministic machine vision algorithm that requires a user-provided length-weight relationship for species of interest and calibration images. 3. Results from validation studies show that lengths and weights can be estimated with high accuracy and precision for round goby ( Neogobius melanostomus ) captured in distorted action camera images, and from large-bodied lake trout ( Salvelinus namaycush ) imaged with a machine vision camera. The real-world utility of the approach is demonstrated in a case study estimating round goby abundances and size frequencies along a 10.7-km transect surveyed with an autonomous underwater vehicle in Lake Michigan, USA. 4. Our validation studies demonstrate that the approach estimates benthic and benthopelagic fish lengths and weights with little bias and good accuracy and precision for species with much different body shapes and sizes. The method is applicable to data collected using a variety of nadir imaging approaches with widespread applications to fisheries monitoring and quantification of any species or object for which nadir images and working distances between the camera and feature of interest are available.

Illinois, Indiana, Michigan, Wisconsin

Analysis and review of fishery-dependent data for Hawaiian nearshore noncommercial fisheries

Noncommercial, shore-based fisheries provide economic, social, and cultural services to communities throughout the Hawaiian Islands. The State of Hawai‘i Department of Land and Natural Resources (DLNR), Division of Aquatic Resources (DAR) routinely conducts surveys to monitor noncommercial fisheries such that estimates of fishing effort and catch by gear type can be generated and used to implement more sustainable management practices. DAR executes both the Hawai‘i Marine Recreational Fishery Survey (HMRFS), a nationally standardized survey that focuses on intercepting fishers at access points (i.e., boat ramps) across the main Hawaiian Islands, and a set of roving creel surveys on O‘ahu, Maui Nui, and Kaua‘i that observe and intercept fishers at locations along the shoreline outside of those targeted by HMRFS. The latter set of creel surveys were designed to complement HMRFS by expanding its geographic coverage and thus providing a more representative picture of noncommercial fishing in Hawai‘i. Sustainable management priorities set by DAR rely on the availability of statewide, fishery- dependent data. Thus, we collate information from island-based roving creel surveys into a cohesive Statewide Creel Survey Database. Further, we provide preliminary analyses and describe ways that surveys could be streamlined to improve future data collection, analysis, and utility. In so doing, we synthesize the most detailed information to-date about noncommercial shore-based fisheries of Hawai‘i. The unprecedented spatial and temporal coverage of DAR’s dataset reveals the value of their survey efforts over the last decade to address fishery management needs. Our primary objectives, results, and conclusions are summarized below: 1) Integrate DAR roving creel survey data from different islands into a single Statewide Creel Survey Dataset (Chapter II). We describe the collation of creel survey data from O‘ahu, Maui Nui, and Kaua‘i into a statewide dataset. We also offer ways in which these surveys could be streamlined to meet the needs of managers and decision makers. Briefly, these are to create a statewide strategic plan, standardize the execution of standard operating procedures, centralize the creel survey database and associated metadata, and consider using technology that improves the data pipeline, including transitioning from paper-based to electronic systems for data entry and processing. 2) Assess whether the new Statewide Creel Survey Dataset can provide inputs for length-based stock assessments (Chapter III). Only on Maui were interviews conducted with associated catch data. There was reasonably high taxonomic coverage (42 species from 186 interviews with 310 fishers), but low sample sizes for nearly all species precluded the development of length-based stock assessments. We provide summary statistics from the existing data and briefly discuss how technologies could be used to automate analysis of images of noncommercial catch. 3) Analyze the Statewide Creel Survey Dataset for spatial and temporal patterns in fishing effort (Chapter IV): a. Visualizing noncommercial fishing pressure . We found that fishing effort (mean number of fishers observed per survey event at a site) on O'ahu was over three times greater than that recorded during similar surveys conducted on Maui or Kaua'i. We create maps that display the distribution of angling and spearfishing effort around each of the three islands. b. Factors that predict fishing “hotspots” around Maui . Fishing effort on Maui was associated with areas with more wave power and less parking availability. There were half as many fishers in areas with parking lots than in areas with parking on the road shoulder only. c. Changes in fishing effort during the COVID-19 pandemic . There was no change in fishing effort on O‘ahu during the first year of the pandemic, but there was a 20% decline in year 2 and a 33% decline in year 3, both in comparison to pre-pandemic levels. Pre-pandemic creel survey data were unavailable for Maui and Kaua‘i, but fishing effort on these islands also declined as the pandemic progressed at similar or greater rates than those observed on O‘ahu. 4) Quantify potential bias in survey methods by experimentally deriving fisher detection probabilities of shore-based and drone-based surveys (Chapter V): a. Shore-based surveys . We conducted roving creel surveys for four months at three locations around Hilo Bay, designed to emulate and estimate the efficacy of DAR standard operating procedures. There was high agreement between paired observers in counting fishers, leading to near-perfect detection probabilities of both anglers (94%) and spearfishers (97%), but relatively low agreement and detection probabilities of other fishers (throw net, ‘opihi picking, etc.) (52%). b. Drone-based surveys . We used an unmanned aerial vehicle (UAV; operated by DAR staff) to collect imagery of fishers along the Hilo Bay shoreline. We used still images and video clips (with known fishing activity) to build an online survey that was distributed to DAR and HCFRU personnel, asking them to count and categorize resource users as a snorkeler, spearfisher, angler, or other fisher. Only 40.0% of the responses correctly counted and categorized resource users in the image. Anglers were correctly identified and enumerated in 90.0% of the responses, but the correct response rates of the other three user categories ranged from 67.8% – 79.4%. Snorkelers and anglers tended to be undercounted while spearfishers and other fishers were overcounted. 5) Review the potential for incorporating emerging technologies that will improve, augment, and evolve creel survey data collection, especially for spearfishing (Chapter VI). Within the context of monitoring shore-based noncommercial fishing, we review the use of electronic data entry/processing systems with geospatial and image capabilities, field cameras, drones, smart buoys, citizen science apps, data mining social media, artificial intelligence and machine learning. We highlight several of the challenges and considerations when implementing these technologies into creel surveys and provide a synthesis of options that could be used to better estimate spearfishing. The general conclusion of this assessment is that the DAR roving creel survey program is collecting valuable data that supplement the existing HMRFS efforts. However, there are a number of areas that could be improved to make these efforts a more effective tool for decision-making processes in resource management and conservation: 1) Establishment of clear statewide and island objectives for the Statewide Creel Survey Dataset. Currently, data collection efforts are focused towards addressing a very broad purpose – supplementing the HMRFS data collection efforts. However, the results of the preliminary analyses conducted as part of this project suggest that the data could be used to address other areas of need if these objectives were clearly defined. Further, the design of the creel survey would benefit from greater standardization of survey protocols between islands and an effort to define a) the acceptable margins of error associated with the estimates generated by these data and b) the minimum level of change that the surveys would need to detect to be useful to managers. 2) Centralization of data entry, data quality assessment, and data accessibility. Currently, each DAR office manages data entry, checks the data for errors, and is responsible for managing and storing the data. Instituting a centralized data entry system, particularly an online database that can receive survey data from tablets or smartphones running a standardized data collection application would improve efficiency, reduce data entry errors, and accelerate the availability of data to managers. A substantial amount of time and effort from the project described in this report was devoted to checking the dataset for errors. The development and application of data quality assurance protocols would ensure that the data are reliable and available in a timely fashion to support management decisions. 3) Address lingering questions regarding the efficacy of current survey protocols to capture and characterize the spearfishing component of the noncommercial fishery. The results presented in the report suggest that the current creel survey protocols do a good job detecting spearfishers when present but are not capturing sufficient data about their catch or total effort. There are also questions remaining as to whether the survey times and sites are sufficiently capturing the behavior of spearfishers in Hawai‘i. A more thorough assessment – whether through additional research, alteration of survey design, or review of data by representatives of the spearfishing community – would provide insight on how to use the Statewide Creel Survey Database to inform management of spearfishing. 4) Investigate the integration of technological advancements into the creel survey methods. As priorities and needs are developed and formalized, it would be valuable to consider how various technological advancements might enhance and streamline data collection or open new avenues of inquiry.

Hawaii

Assessing the potential for spectrally based remote sensing of salmon spawning locations

Remote sensing tools are increasingly used for quantitative mapping of fluvial habitats, yet few techniques exist for continuous sampling of aquatic organisms, such as spawning salmonids. This study assessed the potential for spectrally based remote sensing of salmon spawning locations (i.e., redds) using data acquired from unmanned aircraft systems (UAS) along a large, gravel‐bed river. We developed a novel, semi‐automated approach for detecting salmon redds by applying machine learning classification and object detection techniques to UAS‐based imagery. We found that both true colour (RGB) and hyperspectral imagery could be used to identify salmon redds, though with varying degrees of accuracy. Redds were mapped with accuracies of ~0.75 from RGB imagery using logistic regression and support vector machines (SVM) classification algorithms, but this type of data could not be used to identify redds using Object‐based Image Analysis (OBIA). The hyperspectral imagery was more useful for mapping salmon redds, with accuracies greater than 0.9 for both logistic regression and SVM classifiers; OBIA of the hyperspectral data resulted in redd detection accuracies up to 0.86. The hyperspectral imagery also yielded complementary physical habitat information including water depth and substrate composition, which we quantified on the basis of a spectrally based chlorophyll absorption ratio. Overall, the hyperspectral imagery more effectively identified salmon spawning locations than RGB images and was more conducive to the classification approaches we evaluated. Each type of remotely sensed data had advantages and limitations, which are important for potential users to understand when incorporating UAS‐based data collection into river ecosystem studies.

California