Geology ReportsSearch

SEARCH · Geology Reports

Results for “Machine Learning with Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

112 records · Page 7Linked to original sources

Wave runup and total water level observations from time series imagery at several sites with varying nearshore morphologies

Coastal imaging systems have been developed to measure wave runup and total water level (TWL) at the shoreline, which is a key metric for assessing coastal flooding and erosion. However, extracting quantitative measurements from coastal images has typically been done through the laborious task of hand-digitization of wave runup timestacks. Timestacks are images created by sampling a cross-shore array of pixels from an image through time as waves propagate towards and run up a beach. We utilize over 7000 hand-digitized timestacks from six diverse locations to train and validate machine learning models to automate the process of TWL extraction. Using these data, we evaluate two deep learning model architectures for the task of runup detection. One is based on a fully convolutional architecture trained from scratch, and the other is a transformer-based architecture trained using transfer learning. The deep learning models provide a probability of each pixel being either wet or dry. When contoured at the 50% level (equal chance of being wet or dry), the deep learning models more accurately identified TWL maxima than minima at all sites. This resulted in accurate predictions of 2% exceedance runup, but under predictions of significant swash and over predictions of wave setup. Improved agreement with the complete TWL time series was obtained through post-processing by utilizing the wet/dry probability of each pixel to weight the contouring toward lower dryness probabilities for runup minima (maxima agreed well with observations without tuning). Overall, a transformer-based model using transfer learning provided the best agreement with wave runup statistics, including a) the 2% exceedance runup, b) significant swash, and c) wave setup at the shoreline. For a random subset of images, the model was found to be within the uncertainty range of hand-digitization. The relative success of the transfer learning model suggests that fine-tuning a large model has advantages compared to training a smaller model from scratch. Models provide per-pixel probabilistic estimates in less than 10 s per timestack on a single computational unit, versus the more than 5 min required for hand-digitization. The model is therefore well-suited for near real-time applications, allowing for the development of early warning systems for difficult to forecast events. Real-time wave runup and total water level observations can also be incorporated into coastal hazards forecasts for data assimilation and continual model validation and improvement.

Coastal Engineering

Where the wild things are: Predicting hotspots of seabird aggregations in the California Current System

Marine Protected Areas (MPAs) provide an important tool for conservation of marine ecosystems. To be most effective, these areas should be strategically located in a manner that supports ecosystem function. To inform marine spatial planning and support strategic establishment of MPAs within the California Current System, we identified areas predicted to support multispecies aggregations of seabirds (“hotspots”). We developed habitat‐association models for 16 species using information from at‐sea observations collected over an 11‐year period (1997–2008), bathymetric data, and remotely sensed oceanographic data for an area from north of Vancouver Island, Canada, to the USA/Mexico border and seaward 600 km from the coast. This approach enabled us to predict distribution and abundance of seabirds even in areas of few or no surveys. We developed single‐species predictive models using a machine‐learning algorithm: bagged decision trees. Single‐species predictions were then combined to identify potential hotspots of seabird aggregation, using three criteria: (1) overall abundance among species, (2) importance of specific areas (“core areas”) to individual species, and (3) predicted persistence of hotspots across years. Model predictions were applied to the entire California Current for four seasons (represented by February, May, July, and October) in each of 11 years. Overall, bathymetric variables were often important predictive variables, whereas oceanographic variables derived from remotely sensed data were generally less important. Predicted hotspots often aligned with currently protected areas (e.g., National Marine Sanctuaries), but we also identified potential hotspots in Northern California/Southern Oregon (from Cape Mendocino to Heceta Bank), Southern California (adjacent to the Channel Islands), and adjacent to Vancouver Island, British Columbia, that are not currently included in protected areas. Prioritization and identification of multispecies hotspots will depend on which group of species is of highest management priority. Modeling hotspots at a broad spatial scale can contribute to MPA site selection, particularly if complemented by fine‐scale information for focal areas.

California

Real-time invasive sea lamprey detection using machine learning classifier models on embedded systems

Invasive sea lamprey ( Petromyzon marinus ) has historically inflicted considerable economic and ecological damage in the Great Lakes and continues to be a major threat. Accurately monitoring sea lampreys are critical to enabling the deployment of more targeted and effective control measures to minimize the impact associated with this species. This paper presents the first stand-alone system for real-time detection of sea lamprey attachment on underwater surfaces through the use of classifier models deployed on a microcontroller system. A range of low-complexity models was explored: single-layer artificial neural networks, logistic regression, Gaussian Naive-Bayes, decision trees, random forest, and Scalable, Efficient, and Fast classifieR (SEFR). Threshold models tuned using a multi-objective optimization formulation were also considered. Classifier models were trained with a dataset generated through live animal testing and presented accuracies between 80 and 86%. The models were deployed on an Arduino microcontroller platform and compared in classification accuracy, detection performance, time complexity, and memory size using real-time detection testing. Classification accuracies between 65 and 75% were observed during validation. Models demonstrated good capture rates for lamprey attachments (63–85%), and average detection delays ranging from 9 to 36 s. A video demonstrating the operation of the system during a real-time validation test is also included in this work. While there is room for improving the accuracy of the system, this research presents the first step toward an electronic sea lamprey monitoring system that can provide a detailed view of sea lamprey activity enhancing control and conservation efforts across its entire range.

Neural Computing and Applications

A big data–model integration approach for predicting epizootics and population recovery in a keystone species

Infectious diseases pose a significant threat to global health and biodiversity. Yet, predicting the spatiotemporal dynamics of wildlife epizootics remains challenging. Disease outbreaks result from complex nonlinear interactions among a large collection of variables that rarely adhere to the assumptions of parametric regression modeling. We adopted a nonparametric machine learning approach to model wildlife epizootics and population recovery, using the disease system of colonial black-tailed prairie dogs (BTPD, Cynomys ludovicianus ) and sylvatic plague as an example. We synthesized colony data between 2001 and 2020 from eight USDA Forest Service National Grasslands across the range of BTPDs in central North America. We then modeled extinctions due to plague and colony recovery of BTPDs in relation to complex interactions among climate, topoedaphic variables, colony characteristics, and disease history. Extinctions due to plague occurred more frequently when BTPD colonies were spatially clustered, in closer proximity to colonies decimated by plague during the previous year, following cooler than average temperatures the previous summer, and when wetter winter/springs were preceded by drier summers/falls. Rigorous cross-validations and spatial predictions indicated that our final models predicted plague outbreaks and colony recovery in BTPD with high accuracy (e.g., AUC generally >0.80). Thus, these spatially explicit models can reliably predict the spatial and temporal dynamics of wildlife epizootics and subsequent population recovery in a highly complex host–pathogen system. Our models can be used to support strategic management planning (e.g., plague mitigation) to optimize benefits of this keystone species to associated wildlife communities and ecosystem functioning. This optimization can reduce conflicts among different landowners and resource managers, as well as economic losses to the ranching industry. More broadly, our big data–model integration approach provides a general framework for spatially explicit forecasting of disease-induced population fluctuations for use in natural resource management decision-making.

Arizona, Colorado, Kansas, Montana, Nebraska, New