Geology ReportsSearch

SEARCH · Geology Reports

Results for “Neural Computing and Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

14 recordsLinked to original sources

Real-time invasive sea lamprey detection using machine learning classifier models on embedded systems

Invasive sea lamprey ( Petromyzon marinus ) has historically inflicted considerable economic and ecological damage in the Great Lakes and continues to be a major threat. Accurately monitoring sea lampreys are critical to enabling the deployment of more targeted and effective control measures to minimize the impact associated with this species. This paper presents the first stand-alone system for real-time detection of sea lamprey attachment on underwater surfaces through the use of classifier models deployed on a microcontroller system. A range of low-complexity models was explored: single-layer artificial neural networks, logistic regression, Gaussian Naive-Bayes, decision trees, random forest, and Scalable, Efficient, and Fast classifieR (SEFR). Threshold models tuned using a multi-objective optimization formulation were also considered. Classifier models were trained with a dataset generated through live animal testing and presented accuracies between 80 and 86%. The models were deployed on an Arduino microcontroller platform and compared in classification accuracy, detection performance, time complexity, and memory size using real-time detection testing. Classification accuracies between 65 and 75% were observed during validation. Models demonstrated good capture rates for lamprey attachments (63–85%), and average detection delays ranging from 9 to 36 s. A video demonstrating the operation of the system during a real-time validation test is also included in this work. While there is room for improving the accuracy of the system, this research presents the first step toward an electronic sea lamprey monitoring system that can provide a detailed view of sea lamprey activity enhancing control and conservation efforts across its entire range.

Neural Computing and Applications

NABat ML: Utilizing deep learning to enable crowdsourced development of automated, scalable solutions for documenting North American bat populations

Bats play crucial ecological roles and provide valuable ecosystem services, yet many populations face serious threats from various ecological disturbances. The North American Bat Monitoring Program (NABat) aims to use its technology infrastructure to assess status and trends of bat populations, while developing innovative and community-driven conservation solutions. Here, we present NABat ML , an automated machine-learning algorithm that improves the scalability and scientific transparency of NABat acoustic monitoring. This model combines signal processing techniques and convolutional neural networks (CNNs) to detect and classify recorded bat echolocation calls. We developed our CNN model with internet-based computing resources (‘cloud environment’), and trained it on >600,000 spectrogram images. We also incorporated species range maps to improve the robustness and accuracy of the model for future ‘unseen’ data. We evaluated model performance using a comprehensive, independent, holdout dataset. NABat ML successfully distinguished 31 classes (30 species and a noise class) with overall weighted-average accuracy and precision rates of 92%, and ≥90% classification accuracy for 19 of the bat species. Using a single cloud-environment computing instance, the entire model training process took <16 h. Synthesis and applications . Our convolutional neural network (CNN)-based model, NABat ML , classifies 30 North American bat species using their recorded echolocation calls with an overall accuracy of 92%. In addition to providing highly accurate species-level classification, NABat ML and its outputs are compatible with Bayesian and other statistical techniques for measuring uncertainty in classification. Our model is open-source and reproducible, enabling future implementations as software on end-user devices and cloud-based web applications. These qualities make NABat ML highly suitable for applications ranging from grassroots community science initiatives to big-data methods developed and implemented by researchers and professional practitioners. We believe the transparency and accessibility of NABat ML will encourage broad-scale participation in bat monitoring, and enable development of innovative solutions needed to conserve North American bat species.

Journal of Applied Ecology

Forecasting water levels using machine (deep) learning to complement numerical modelling in the southern Everglades, USA

Water level is an important guide for water resource management and wetland ecosystems, defining one of the most basic processes in hydrology. This research seeks to investigate the possibility of complementing numerical modeling with a Machine Learning (ML) model to forecast daily water levels in the southern Everglades in Florida, USA. An exact analytical solution to water level may not be possible, but using the computational methods afforded by ML, the traditional numerical techniques may be enhanced to generate more robust, scalable predictions. Five locations were chosen for application of the Time-Delayed Neural Network (TDNN) and Long-Short Term Memory Recurrent Neural Network (LSTM-RNN) ML models, which were built to estimate water level with 1, 2, 3, 7 and 10 day forecasts using a simulation step of 1 day. The results showed that rainfall forecasts from weather models could improve water-level forecasts if the accuracy and performance of the weather models can be improved. The ML models presented here improve water-level predictions from a historical hydrologic model for a 24 hour forecast horizon.

Florida

Hybrid modeling of spatial continuity for application to numerical inverse problems

A novel two-step modeling approach is presented to obtain optimal starting values and geostatistical constraints for numerical inverse problems otherwise characterized by spatially-limited field data. First, a type of unsupervised neural network, called the self-organizing map (SOM), is trained to recognize nonlinear relations among environmental variables (covariates) occurring at various scales. The values of these variables are then estimated at random locations across the model domain by iterative minimization of SOM topographic error vectors. Cross-validation is used to ensure unbiasedness and compute prediction uncertainty for select subsets of the data. Second, analytical functions are fit to experimental variograms derived from original plus resampled SOM estimates producing model variograms. Sequential Gaussian simulation is used to evaluate spatial uncertainty associated with the analytical functions and probable range for constraining variables. The hybrid modeling of spatial continuity is demonstrated using spatially-limited hydrologic measurements at different scales in Brazil: (1) physical soil properties (sand, silt, clay, hydraulic conductivity) in the 42 km 2 Vargem de Caldas basin; (2) well yield and electrical conductivity of groundwater in the 132 km 2 fractured crystalline aquifer; and (3) specific capacity, hydraulic head, and major ions in a 100,000 km 2 transboundary fractured-basalt aquifer. These results illustrate the benefits of exploiting nonlinear relations among sparse and disparate data sets for modeling spatial continuity, but the actual application of these spatial data to improve numerical inverse modeling requires testing.

Environmental Modelling and Software

Physics-guided recurrent neural networks for predicting lake water temperature

This chapter presents a physics-guided recurrent neural network model (PGRNN) for predicting water temperature in lake systems. Standard machine learning (ML) methods, especially deep learning models, often require a large amount of labeled training samples, which are often not available in scientific problems due to the substantial human labor and material costs associated with data collection. ML models have found tremendous success in several commercial applications, e.g., computer vision and natural language processing. The chapter presents PGRNN as a general framework for modeling physical processes in engineering and environmental systems. The proposed PGRNN explicitly incorporates physical laws such as energy conservation or mass conservation. In particular, researchers started pursing this direction by using residual modeling, where an ML model is learned to predict the errors, or residuals, made by a physics-based model. Advanced ML models, especially deep learning models, often require a large amount of training data for tuning model parameters.

Book chapter

Evaluating the sources of water to wells: Three techniques for metamodeling of a groundwater flow model

For decision support, the insights and predictive power of numerical process models can be hampered by insufficient expertise and computational resources required to evaluate system response to new stresses. An alternative is to emulate the process model with a statistical &ldquo;metamodel.&rdquo; Built on a dataset of collocated numerical model input and output, a groundwater flow model was emulated using a Bayesian Network, an Artificial neural network, and a Gradient Boosted Regression Tree. The response of interest was surface water depletion expressed as the source of water-to-wells. The results have application for managing allocation of groundwater. Each technique was tuned using cross validation and further evaluated using a held-out dataset. A numerical MODFLOW-USG model of the Lake Michigan Basin, USA, was used for the evaluation. The performance and interpretability of each technique was compared pointing to advantages of each technique. The metamodel can extend to unmodeled areas.

Illinois, Indiana, Michigan, Ohio, Wisconsin

Statistical approach to neural network imaging of karst systems in 3D seismic reflection data

The current lack of a robust, standardized technique for geophysical mapping of karst systems can be attributed to both the complexity of the environment and prior technological limitations. Abrupt lateral variations in physical properties that are inherent to karst systems generate significant geophysical noise, challenging conventional seismic signal processing and interpretation. Modern application of neural networks to multi-attribute seismic interpretation now provide a semiautomated method for identifying and leveraging the nonlinear relationships exhibited among seismic attributes. The ambiguity generally associated with designing neural networks for seismic object detection can be reduced via statistical analysis of the extracted attribute data. A data-driven approach to selecting the appropriate set of input seismic attributes, as well as the locations and minimum number of training examples, provides a more objective and computationally efficient method for identifying karst systems using reflection seismology. This statistically optimized neural network technique is thoroughly demonstrated using three-dimensional seismic reflection data collected from the southeastern portion of the Florida carbonate platform. Several dimensionality reduction methods are applied and the resulting karst probability models are evaluated relative to one another based on both quantitative and qualitative criteria. Comparing the preferred model, using quadratic discriminant analysis, to previously available seismic object detection workflows demonstrates the karst-specific nature of the tool. Results suggest that the karst multi-attribute workflow presented is capable of approximating the structural boundaries of karst systems with more accuracy and efficiency than a human counterpart or previously presented seismic interpretation schemes. This objective technique, using solely three-dimensional seismic reflection data, likely represents the most practical approach to mapping karst systems for subsequent hydrogeological modeling.

Interpretation

Projecting impacts of climate change on water availability using artificial neural network techniques

Lago Loíza reservoir in east-central Puerto Rico is one of the primary sources of public water supply for the San Juan metropolitan area. To evaluate and predict the Lago Loíza water budget, an artificial neural network (ANN) technique is trained to predict river inflows. A method is developed to combine ANN-predicted daily flows with ANN-predicted 30-day cumulative flows to improve flow estimates. The ANN application trains well for representing 2007–2012 and the drier 1994–1997 periods. Rainfall data downscaled from global circulation model (GCM) simulations are used to predict 2050–2055 conditions. Evapotranspiration is estimated with the Hargreaves equation using minimum and maximum air temperatures from the downscaled GCM data. These simulated 2050–2055 river flows are input to a water budget formulation for the Lago Loíza reservoir for comparison with 2007–2012. The ANN scenarios require far less computational effort than a numerical model application, yet produce results with sufficient accuracy to evaluate and compare hydrologic scenarios. This hydrologic tool will be useful for future evaluations of the Lago Loíza reservoir and water supply to the San Juan metropolitan area.

Puerto Rico

A review of supervised learning methods for classifying animal behavioural states from environmental features

Accurately predicting behavioural modes of animals in response to environmental features is important for ecology and conservation. Supervised learning (SL) methods are increasingly common in animal movement ecology for classifying behavioural modes. However, few examples exist of applying SL to classify polytomous animal behaviour from environmental features especially in the context of millions of animal observations. We review SL methods (weighted k -nearest neighbours; neural nets; random forests; and boosted classification trees with XGBoost) for classifying polytomous animal behaviour from environmental predictors. We also describe tuning parameter selection and assessment strategies, approaches for visualizing relationships between predictors and class outputs, and computational considerations. We demonstrate these methods by predicting three categories of risk to bald eagles from colliding with wind turbines using, as predictors, 12 environmental state features associated with 1.7 million GPS telemetry data points from 57 eagles. Of the SL methods we considered, XGBoost yielded the most accurate model with 86.2% classification accuracy and pairwise-averaged area under the ROC curve of 90.6. Computational time of XGBoost scaled better to large data than any other SL method. We also show how SHAP values integrated in the R package ( xgboost ) facilitate investigation of variable relationships and importance. For big data applications, XGBoost appears to provide superior classification accuracy and computational efficiency. Our results suggest XGBoost should be considered as an early modelling option in situations where the intent is to classify millions of animal behaviour observations from environmental predictors and to understand relationships between those predictors and movement behaviours. We also offer a tutorial to assist researchers in implementing this method.

Methods in Ecology and Evolution

Weakly supervised spatial deep learning for Earth image segmentation based on imperfect polyline labels

In recent years, deep learning has achieved tremendous success in image segmentation for computer vision applications. The performance of these models heavily relies on the availability of large-scale high-quality training labels (e.g., PASCAL VOC 2012). Unfortunately, such large-scale high-quality training data are often unavailable in many real-world spatial or spatiotemporal problems in earth science and remote sensing (e.g., mapping the nationwide river streams for water resource management). Although extensive efforts have been made to reduce the reliance on labeled data (e.g., semi-supervised or unsupervised learning, few-shot learning), the complex nature of geographic data such as spatial heterogeneity still requires sufficient training labels when transferring a pre-trained model from one region to another. On the other hand, it is often much easier to collect lower-quality training labels with imperfect alignment with earth imagery pixels (e.g., through interpreting coarse imagery by non-expert volunteers). However, directly training a deep neural network on imperfect labels with geometric annotation errors could significantly impact model performance. Existing research that overcomes imperfect training labels either focuses on errors in label class semantics or characterizes label location errors at the pixel level. These methods do not fully incorporate the geometric properties of label location errors in the vector representation. To fill the gap, this article proposes a weakly supervised learning framework to simultaneously update deep learning model parameters and infer hidden true vector label locations. Specifically, we model label location errors in the vector representation to partially reserve geometric properties (e.g., spatial contiguity within line segments). Evaluations on real-world datasets in the National Hydrography Dataset (NHD) refinement application illustrate that the proposed framework outperforms baseline methods in classification accuracy.

ACM Transactions on Intelligent Systems and Techno

Improving the accessibility and transferability of machine learning algorithms for identification of animals in camera trap images: MLWIC2

Motion‐activated wildlife cameras (or “camera traps”) are frequently used to remotely and noninvasively observe animals. The vast number of images collected from camera trap projects has prompted some biologists to employ machine learning algorithms to automatically recognize species in these images, or at least filter‐out images that do not contain animals. These approaches are often limited by model transferability, as a model trained to recognize species from one location might not work as well for the same species in different locations. Furthermore, these methods often require advanced computational skills, making them inaccessible to many biologists. We used 3 million camera trap images from 18 studies in 10 states across the United States of America to train two deep neural networks, one that recognizes 58 species, the “species model,” and one that determines if an image is empty or if it contains an animal, the “empty‐animal model.” Our species model and empty‐animal model had accuracies of 96.8% and 97.3%, respectively. Furthermore, the models performed well on some out‐of‐sample datasets, as the species model had 91% accuracy on species from Canada (accuracy range 36%–91% across all out‐of‐sample datasets) and the empty‐animal model achieved an accuracy of 91%–94% on out‐of‐sample datasets from different continents. Our software addresses some of the limitations of using machine learning to classify images from camera traps. By including many species from several locations, our species model is potentially applicable to many camera trap studies in North America. We also found that our empty‐animal model can facilitate removal of images without animals globally. We provide the trained models in an R package (MLWIC2: Machine Learning for Wildlife Image Classification in R), which contains Shiny Applications that allow scientists with minimal programming experience to use trained models and train new models in six neural network architectures with varying depths.

Ecology and Evolution

Automatic identification and quantification of volcanic hotspots in Alaska using HotLINK: The hotspot learning and identification network

An increase in volcanic thermal emissions can indicate subsurface and surface processes that precede, or coincide with, volcanic eruptions. Space-borne infrared sensors can detect hotspots—defined here as localized volcanic thermal emissions—in near-real-time. However, automatic hotspot detection systems are needed to efficiently analyze the large quantities of data produced. While hotspots have been automatically detected for over 20 years with simple thresholding algorithms, new computer vision technologies, such as convolutional neural networks (CNNs), can enable improved detection capabilities. Here we introduce HotLINK: the Hotspot Learning and Identification Network, a CNN trained to detect hotspots with a dataset of −3,800 satellite-based, Visible Infrared Imaging Radiometer Suite (VIIRS) images from Mount Veniaminof and Mount Cleveland volcanoes, Alaska. We find that our model achieves an accuracy of 96% (F1-score 0.92) when evaluated on −1,700 unseen images from the same volcanoes, and 95% (F1-score 0.67) when evaluated on −3,000 images from six additional Alaska volcanoes (Augustine Volcano, Bogoslof Island, Okmok Caldera, Pavlof Volcano, Redoubt Volcano, Shishaldin Volcano). In comparison with an existing threshold-based hotspot detection algorithm, MIROVA (Coppola et al., Geological Society, London, Special Publications, 2016, 426, 181–205), our model detects 22% more hotspots and produces 12% fewer false positives. Additional testing on −700 labeled Moderate Resolution Imaging Spectroradiometer (MODIS) images from Mount Veniaminof demonstrates that our model is applicable to this sensor’s data as well, achieving an accuracy of 98% (F1-score 0.95). We apply HotLINK to 10 years of VIIRS data and 22 years of MODIS data for the eight aforementioned Alaska volcanoes and calculate the radiative power of detected hotspots. From these time series we find that HotLINK accurately characterizes background and eruptive periods, similar to MIROVA, but also detects more subtle warming signals, potentially related to volcanic unrest. We identify three advantages to our model over its predecessors: 1) the ability to detect more subtle volcanic hotspots and produce fewer false positives, especially in daytime images; 2) probabilistic predictions provide a measure of detection confidence; and 3) its transferability, i.e., the successful application to multiple sensors and multiple volcanoes without the need for threshold tuning, suggesting the potential for global application.

Alaska

SlideDetect: Spatio-temporal landslide detection using a three-dimensional convolutional neural network

Landslides pose a serious and ongoing threat to both human lives and infrastructure worldwide; therefore, it is of interest to predict where and when landslides are likely to occur. Advances in machine learning techniques have spurred numerous studies aimed at estimating relative landslide propensity, but are limited to spatial (as opposed to temporal) prediction due to the sparsity of landslide timing data. We address this data gap by training SlideDetect, a 3-dimensional convolutional neural network (3D CNN), to identify landslides based on their spatial and temporal occurrence within multitemporal image stacks. We use an inventory of landsides triggered by the 2018 Hokkaido earthquake and two years of monthly composite optical imagery spanning this event. The model can identify not only landslide location but also landslide date with an area under the precision-recall curve (PR-AUC) of 0.84. We further present a new standard for presenting PR curve results that explicitly compares model performance at different confidence thresholds, allowing for clearer model evaluation and comparison. Our new approach to constraining landslide timing paired with this more consistent and objective method for evaluating model performance shows considerable promise, and with further application and testing, SlideDetect could enhance the data availability and tools needed to advance landslide hazard and risk assessments.

JGR Machine Learning and Computation

Methods to evaluate and improve the modeling of rupture directivity in assessment of seismic hazard

In recent years, there have been several advancements related to the modelling of near-source effects of earthquake rupture on strong ground shaking, leading to an improved characterization of ground motions and resulting seismic hazard. Some of these modifications have stemmed from physics-based numerical modelling of the earthquake rupture process, using physics-based dynamic rupture simulations. These contributions have led to a better understanding of how fault rupture characteristics, geometry, and the style of faulting can interact with the hypocenter-dependence on the path from source to site that may ultimately guide the development of seismic directivity models. Moving forward, the application of modern techniques can be used to incorporate these source characteristics and near-fault ground motion behavior that contribute to the azimuthally varying effects that result in rupture directivity. One example is the application of machine learning methods to support more automated integration of new predictor variables in model development and open more evaluation opportunities to access residuals. Here, we utilize several techniques to take advantage of the plethora of synthetic data and its ability to supplement preexisting trends observed in data. We showcase two examples of how models can be either developed, expanded upon, or constrained using artificial neural network model (ANNs). We evaluate the performance of the ANN with existing methods, comparing misfit, potential limitations, and ability to continue to improve upon these methods in the future. One approach uses a set of simulations with corresponding synthetic ground motions from the Southern California Earthquake Center (SCEC) CyberShake study to develop a ground motion model adapted to incorporate seismic directivity information using an ANN. This large database (TBs) enables us to train the model to capture magnitude, period, and distance variations and how these parameters relate to amplification from hypocenters located along finite-faults. In some cases, there is reduced misfit from better representing source features that aren’t included in base ground motion models that neglect hypocenter location (e.g. azimuthal variation, source-to-site terms). Another ANN method uses a shallow-layered neural network model to better fit a hypocenter-independent model. This method adjusts the median and aleatory variability to account for the averaged impact of various hypocenter distributions to fit the underlying directivity adjustment model. This method serves as a template to apply to other directivity models, improving computational efficiency and more readily enabling integration in hazard codes.

California