Geology ReportsSearch

SEARCH · Geology Reports

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Selection and quality assessment of Landsat data for the North American forest dynamics forest history maps of the US

Using the NASA Earth Exchange platform, the North American Forest Dynamics (NAFD) project mapped forest history wall-to-wall, annually for the contiguous US (1986–2010) using the Vegetation Change Tracker algorithm. As with any effort to identify real changes in remotely sensed time-series, data gaps, shifts in seasonality, misregistration, inconsistent radiometry and cloud contamination can be sources of error. We discuss the NAFD image selection and processing stream (NISPS) that was designed to minimize these sources of error. The NISPS image quality assessments highlighted issues with the Landsat archive and metadata including inadequate georegistration, unreliability of the pre-2009 L5 cloud cover assessments algorithm, missing growing-season imagery and paucity of clear views. Assessment maps of Landsat 5–7 image quantities and qualities are presented that offer novel perspectives on the growing-season archive considered for this study. Over 150,000+ Landsat images were considered for the NAFD project. Optimally, one high quality cloud-free image in each year or a total of 12,152 images would be used. However, to accommodate data gaps and cloud/shadow contamination 23,338 images were needed. In 220 specific path-row image years no acceptable images were found resulting in data gaps in the annual national map products.

International Journal of Digital Earth

Applications of satellite ocean color sensors for monitoring and predicting harmful algal blooms

The new satellite ocean color sensors offer a means of detecting and monitoring algal blooms in the ocean and coastal zone. Beginning with SeaWiFS (Sea Wide Field-of-view Sensor) in September 1997, these sensors provide coverage every 1 to 2 days with 1-km pixel view at nadir. Atmospheric correction algorithms designed for the coastal zone combined with regional chlorophyll algorithms can provide good and reproducible estimates of chlorophyll, providing the means of monitoring various algal blooms. Harmful algal blooms (HABs) caused by Karenia brevis in the Gulf of Mexico are particularly amenable to remote observation. The Gulf of Mexico has relatively clear water and K. brevis , in bloom conditions, tends to produce a major portion of the phytoplankton biomass. A monitoring program has begun in the Gulf of Mexico that integrates field data from state monitoring programs with satellite imagery, providing an improved capability for the monitoring of K. brevis blooms.

Gulf of Mexico

Temporal expansion of annual crop classification layers for the CONUS using the C5 decision tree classifier

Crop cover maps have become widely used in a range of research applications. Multiple crop cover maps have been developed to suite particular research interests. The National Agricultural Statistics Service (NASS) Cropland Data Layers (CDL) are a series of commonly used crop cover maps for the conterminous United States (CONUS) that span from 2008 to 2013. In this investigation, we sought to contribute to the availability of consistent CONUS crop cover maps by extending temporal coverage of the NASS CDL archive back eight additional years to 2000 by creating annual NASS CDL-like crop cover maps derived from a classification tree model algorithm. We used over 11 million records to train a classification tree algorithm and develop a crop classification model (CCM). The model was used to create crop cover maps for the CONUS for years 2000–2013 at 250 m spatial resolution. The CCM and the maps for years 2008–2013 were assessed for accuracy relative to resampled NASS CDLs. The CCM performed well against a withheld test data set with a model prediction accuracy of over 90%. The assessment of the crop cover maps indicated that the model performed well spatially, placing crop cover pixels within their known domains; however, the model did show a bias towards the ‘Other’ crop cover class, which caused frequent misclassifications of pixels around the periphery of large crop cover patch clusters and of pixels that form small, sparsely dispersed crop cover patches.

Remote Sensing Letters

Evaluating machine learning approaches to identify and predict oil and gas produced water lithium concentrations

Recently, the demand for battery-grade lithium has substantially increased, largely due to electrification of the transportation sector. The search for new lithium sources has turned to produced waters (frequently brines), a large-volume wastewater by-product of oil and gas extraction. Geochemical analysis indicates the presence of varying concentrations of lithium from produced water samples collected across the United States and represented in the U.S. Geological Survey’s National Produced Water Geochemical Database, as well as mixtures of Marcellus Shale produced water included in the Pennsylvania Department of Environmental Protection’s Oil and Gas Well Waste Reports. We first examined whether the geochemical signature of the lithium-bearing produced waters is sufficiently distinct so that machine learning (ML) can be used to correctly classify samples to the formation of origin. The produced water sample data used to assess classification accuracy were from the Marcellus Shale, Utica Shale and Point Pleasant Formation (Utica), and Smackover Formation oil and gas wells. Further, we evaluated the potential for ML to accurately classify Marcellus Shale produced water spatially (i.e., northeast versus southwest Pennsylvania). We then investigated whether ML algorithms applied to a suite of geochemical concentration data (i.e. Ba, Br, Cl, K, Mg, Sr) may be used to predict the lithium concentration of an unknown sample. Finally, we applied an estimated economic lithium grade cutoff of 150 milligrams per liter (mg/l) and assessed the utility of ML to predict whether a produced water sample would fall above or below the grade cutoff based on the suite of geochemical parameters. Four machine learning algorithms—Random Forest (RF), Gradient Boosting Trees (GBT), Extreme Boosting (XGBoost), and Deep Neural Networks (DNN) were assessed. This study successfully demonstrates that all four machine learning methods can precisely and accurately estimate lithium concentrations and geologic formation classification. The products of this study contribute to the growing body of knowledge aimed at expanding the lithium resource base within the United States.

Alabama, Arkansas, Florida, Georgia, Louisiana, Mi

CyberShake-derived ground-motion prediction models for the Los Angeles region with application to earthquake early warning

Real-time applications such as earthquake early warning (EEW) typically use empirical ground-motion prediction equations (GMPEs) along with event magnitude and source-to-site distances to estimate expected shaking levels. In this simplified approach, effects due to finite-fault geometry, directivity and site and basin response are often generalized, which may lead to a significant under- or overestimation of shaking from large earthquakes ( M > 6.5) in some locations. For enhanced site-specific ground-motion predictions considering 3-D wave-propagation effects, we develop support vector regression (SVR) models from the SCEC CyberShake low-frequency (<0.5 Hz) and broad-band (0–10 Hz) data sets. CyberShake encompasses 3-D wave-propagation simulations of >415 000 finite-fault rupture scenarios (6.5 ≤ M ≤ 8.5) for southern California defined in UCERF 2.0. We use CyberShake to demonstrate the application of synthetic waveform data to EEW as a ‘proof of concept’, being aware that these simulations are not yet fully validated and might not appropriately sample the range of rupture uncertainty. Our regression models predict the maximum and the temporal evolution of instrumental intensity (MMI) at 71 selected test sites using only the hypocentre, magnitude and rupture ratio, which characterizes uni- and bilateral rupture propagation. Our regression approach is completely data-driven (where here the CyberShake simulations are considered data) and does not enforce pre-defined functional forms or dependencies among input parameters. The models were established from a subset (∼20 per cent) of CyberShake simulations, but can explain MMI values of all >400 k rupture scenarios with a standard deviation of about 0.4 intensity units. We apply our models to determine threshold magnitudes (and warning times) for various active faults in southern California that earthquakes need to exceed to cause at least ‘moderate’, ‘strong’ or ‘very strong’ shaking in the Los Angeles (LA) basin. These thresholds are used to construct a simple and robust EEW algorithm: to declare a warning, the algorithm only needs to locate the earthquake and to verify that the corresponding magnitude threshold is exceeded. The models predict that a relatively moderate M 6.5–7 earthquake along the Palos Verdes, Newport-Inglewood/Rose Canyon, Elsinore or San Jacinto faults with a rupture propagating towards LA could cause ‘very strong’ to ‘severe’ shaking in the LA basin; however, warning times for these events could exceed 30 s.

California

Accurate recapture identification for genetic mark–recapture studies with error-tolerant likelihood-based match calling and sample clustering

Error-tolerant likelihood-based match calling presents a promising technique to accurately identify recapture events in genetic mark–recapture studies by combining probabilities of latent genotypes and probabilities of observed genotypes, which may contain genotyping errors. Combined with clustering algorithms to group samples into sets of recaptures based upon pairwise match calls, these tools can be used to reconstruct accurate capture histories for mark–recapture modelling. Here, we assess the performance of a recently introduced error-tolerant likelihood-based match-calling model and sample clustering algorithm for genetic mark–recapture studies. We assessed both biallelic (i.e. single nucleotide polymorphisms; SNP) and multiallelic (i.e. microsatellite; MSAT) markers using a combination of simulation analyses and case study data on Pacific walrus ( Odobenus rosmarus divergens ) and fishers ( Pekania pennanti ). A novel two-stage clustering approach is demonstrated for genetic mark–recapture applications. First, repeat captures within a sampling occasion are identified. Subsequently, recaptures across sampling occasions are identified. The likelihood-based matching protocol performed well in simulation trials, demonstrating utility for use in a wide range of genetic mark–recapture studies. Moderately sized SNP (64+) and MSAT (10–15) panels produced accurate match calls for recaptures and accurate non-match calls for samples from closely related individuals in the face of low to moderate genotyping error. Furthermore, matching performance remained stable or increased as the number of genetic markers increased, genotyping error notwithstanding.

Royal Society Open Science

Discovering loose group movement patterns from animal trajectories

The technical advances of positioning technologies enable us to track animal movements at finer spatial and temporal scales, and further help to discover a variety of complex interactive relationships. In this paper, considering the loose gathering characteristics of the real-life groups' members during the movements, we propose two kinds of loose group movement patterns and corresponding discovery algorithms. Firstly, we propose the weakly consistent group movement pattern which allows the gathering of a part of the members and individual temporary leave from the whole during the movements. To tolerate the high dispersion of the group at some moments (i.e. to adapt the discontinuity of the group's gatherings), we further scheme the weakly consistent and continuous group movement pattern. The extensive experimental analysis and comparison with the real and synthetic data shows that the group pattern discovery algorithms proposed in this paper are similar to the the real-life frequent divergences of the members during the movements, can discover more complete memberships, and have considerable performance.

Conference Paper

Web-client based distributed generalization and geoprocessing

Generalization and geoprocessing operations on geospatial information were once the domain of complex software running on high-performance workstations. Currently, these computationally intensive processes are the domain of desktop applications. Recent efforts have been made to move geoprocessing operations server-side in a distributed, web accessible environment. This paper initiates research into portable client-side generalization and geoprocessing operations as part of a larger effort in user-centered design for the US Geological Survey's The National Map. An implementation of the Ramer-Douglas-Peucker (RDP) line simplification algorithm was created in the open source OpenLayers geoweb client. This algorithm implementation was benchmarked using differing data structures and browser platforms. The implementation and results of the benchmarks are discussed in the general context of client-side geoprocessing. (Abstract).

Conference Paper

Comparative analysis of multisensor satellite monitoring of Arctic sea-ice

This report represents comparative analysis of nearly coincident Russian OKEAN-01 polar orbiting satellite data, Special Sensor Microwave Imager (SSM/I) and Advanced Very High Resolution Radiometer (AVHRR) imagery. OKEAN-01 ice concentration algorithms utilize active and passive microwave measurements and a linear mixture model for measured values of the brightness temperature and the radar backscatter. SSM/I and AVHRR ice concentrations were computed with NASA Team algorithm and visible and thermal-infrared wavelength AVHRR data, accordingly

Book

Soft pressure sensor for underwater sea lamprey detection

In this paper, an economical and effective soft pressure sensor for underwater sea lamprey detection is proposed, which consists of an array of piezoresistive elements between two layers of perpendicular copper tape electrodes, forming a passive resistor network. With multiplexers, the apparent resistance corresponding to each pixel of the sensing matrix can be measured directly, where the pixel is identified with the row and the column of the respective electrodes. However, this measured two-point resistance is not equal to the actual cell resistance for that pixel due to the crosstalk effect in the resistor network. Since the cell resistance reflects directly the pressure applied on each pixel, the relationship between the cell resistance and the measured two-point resistance is analyzed for a passive matrix of any size. More importantly, several regularized least-squares algorithms are proposed to reconstruct the cell resistance profile from the two-point resistance measurements, with enhanced robustness of the reconstruction in the presence of measurement noises and modeling errors. The proposed pressure sensor is applied to detect the suction attachment of sea lampreys, a devastating invasive species in the Great Lakes region. Experimental results demonstrate that the pressure sensor can successfully capture the rim profile of the lamprey’s sucking mouth. Moreover, the performance and computational complexity of the reconstruction algorithms with different regularization functions are compared.

IEEE Sensors Journal

Automated soft pressure sensor array-based sea lamprey detection using machine learning

Sea lamprey, a destructive invasive species in the Great Lakes in North America, is among very few fishes that rely on oral suction during migration and spawning. Recently, soft pressure sensors have been proposed to detect the attachment of sea lamprey as part of the monitoring and control effort. However, human decision is still required for the recognition of patterns in the measured signals. In this article, a novel automated soft pressure sensor array-based sea lamprey detection framework is proposed using object detection convolutional neural networks. First, the resistance measurements of the pressure sensor array are converted to mappings of relative change in resistance. These mappings typically show two different types of patterns under lamprey attachment: a high-pressure circular pattern corresponding to the mouth rim compressed against the sensor (“compression” pattern), and a low-pressure blob corresponding to the partial vacuum region of the sucking mouth (“suction” pattern). Three types of object detection algorithms, single-shot detector (SSD), RetinaNet, and YOLOv5s, are applied to the dataset of measurements collected in the presence of sea lamprey attachment, and the comparison of their performance shows that YOLOv5s model achieves the highest mean average precision (mAP) and the fastest inference speed. Furthermore, to improve the accuracy of the prediction model and reduce the false positive (FP) rate due to the sensor’s memory effect, a filter branch with different detection thresholds for the compression and suction patterns, respectively, is added to the original machine-learning algorithm. The trained model is validated and used to automatically detect sea lamprey attachments and locate the suction area on the sensor in real time.

IEEE Sensors Journal

In-flight validation and recovery of water surface temperature with Landsat-5 thermal infrared data using an automated high-altitude lake validation site at Lake Tahoe

The absolute radiometric accuracy of the thermal infrared band (B6) of the Thematic Mapper (TM) instrument on the Landsat-5 (L5) satellite was assessed over a period of approximately four years using data from the Lake Tahoe automated validation site (California-Nevada). The Lake Tahoe site was established in July 1999, and measurements of the skin and bulk temperature have been made approximately every 2 min from four permanently moored buoys since mid-1999. Assessment involved using a radiative transfer model to propagate surface skin temperature measurements made at the time of the L5 overpass to predict the at-sensor radiance. The predicted radiance was then convolved with the L5B6 system response function to obtain the predicted L5B6 radiance, which was then compared with the radiance measured by L5B6. Twenty-four cloud-free scenes acquired between 1999 and 2003 were used in the analysis with scene temperatures ranging between 4/spl deg/C and 22/spl deg/C. The results indicate L5B6 had a radiance bias of 2.5% (1.6/spl deg/C) in late 1999, which gradually decreased to 0.8% (0.5/spl deg/C) in mid-2002. Since that time, the bias has remained positive (predicted minus measured) and between 0.3% (0.2/spl deg/C) and 1.4% (0.9/spl deg/C). The cause for the cold bias (L5 radiances are lower than expected) is unresolved, but likely related to changes in instrument temperature associated with changes in instrument usage. The in situ data were then used to develop algorithms to recover the skin and bulk temperature of the water by regressing the L5B6 radiance and the National Center for Environmental Prediction (NCEP) total column water data to either the skin or bulk temperature. Use of the NCEP data provides an alternative approach to the split-window approach used with instruments that have two thermal infrared bands. The results indicate the surface skin and bulk temperature can be recovered with a standard error of 0.6/spl deg/C. This error is larger than errors obtained with other instruments due, in part, to the calibration bias. L5 provides the only long-duration high spatial resolution thermal infrared measurements of the land surface. If these data are to be used effectively in studies designed to monitor change, it is essential to continue to monitor instrument performance in-flight and develop quantitative algorithms for recovering surface temperature.

IEEE Transactions on Geoscience and Remote Sensing

Impact of spectral resolution on quantifying cyanobacteria in lakes and reservoirs: A machine-learning assessment

Cyanobacterial harmful algal blooms are an increasing threat to coastal and inland waters. These blooms can be detected using optical radiometers due to the presence of phycocyanin (PC) pigments. The spectral resolution of best-available multispectral sensors limits their ability to diagnostically detect PC in the presence of other photosynthetic pigments. To assess the role of spectral resolution in the determination of PC, a large (N = 905) database of colocated in situ radiometric spectra and PC are employed. We first examine the performance of selected widely used machine-learning (ML) models against that of benchmark algorithms for hyperspectral remote sensing reflectance ( R r s ) spectra resampled to the spectral configuration of the Hyperspectral Imager for the Coastal Ocean (HICO) with a full-width at half-maximum (FWHM) of < 6 nm. Results show that the multilayer perceptron (MLP) neural network applied to HICO spectral configurations (median errors < 65%) outperforms other ML models. This model is subsequently applied to R r s spectra resampled to the band configuration of existing satellite instruments and of the one proposed for the next Landsat sensor. These results confirm that employing MLP models to estimate PC from hyperspectral data delivers tangible improvements compared with retrievals from multispectral data and benchmark algorithms (with median errors between ~73% and 126%) and shows promise for developing a globally applicable cyanobacteria measurement approach.

IEEE Transactions in Geoscience and Remote Sensing

Ground settlement monitoring from temporarily persistent scatterers between two SAR acquisitions

We present an improved differential interferometric synthetic aperture radar (DInSAR) analysis method that measures motions of scatterers whose phases are stable between two SAR acquisitions. Such scatterers are referred to as temporarily persistent scatterers (TPS) for simplicity. Unlike the persistent scatterer InSAR (PS-InSAR) method that relies on a time-series of interferograms, the new algorithm needs only one interferogram. TPS are identified based on pixel offsets between two SAR images, and are specially coregistered based on their estimated offsets instead of a global polynomial for the whole image. Phase unwrapping is carried out based on an algorithm for sparse data points. The method is successfully applied to measure the settlement in the Hong Kong Airport area. The buildings surrounded by vegetation were successfully selected as TPS and the tiny deformation signal over the area was detected. ??2009 IEEE.

Conference Paper

Near-term forecasts of stream temperature using deep learning and data assimilation in support of management decisions

Deep learning (DL) models are increasingly used to make accurate hindcasts of management-relevant variables, but they are less commonly used in forecasting applications. Data assimilation (DA) can be used for forecasts to leverage real-time observations, where the difference between model predictions and observations today is used to adjust the model to make better predictions tomorrow. In this use case, we developed a process-guided DL and DA approach to make 7-day probabilistic forecasts of daily maximum water temperature in the Delaware River Basin in support of water management decisions. Our modeling system produced forecasts of daily maximum water temperature with an average root mean squared error (RMSE) from 1.1 to 1.4°C for 1-day-ahead and 1.4 to 1.9°C for 7-day-ahead forecasts across all sites. The DA algorithm marginally improved forecast performance when compared with forecasts produced using the process-guided DL model alone (0%–14% lower RMSE with the DA algorithm). Across all sites and lead times, 65%–82% of observations were within 90% forecast confidence intervals, which allowed managers to anticipate probability of exceedances of ecologically relevant thresholds and aid in decisions about releasing reservoir water downstream. The flexibility of DL models shows promise for forecasting other important environmental variables and aid in decision-making.

Journal of the American Water Resources Associatio

Leveraging high-frequency sensor data and U.S. National Water Model output to forecast turbidity in a drinking water supply basin

As high-frequency sensor networks increasingly enhance data-driven models of water quality, process-based models like the U.S. National Water Model (NWM) are generating accessible forecasts of streamflow at increasingly dense scales. There is now an opportunity to combine these products to construct actionable water quality forecasts. To that end, we couple streamflow forecasts from the NWM to a gradient-boosted decision tree algorithm (LightGBM) trained on 5+ years of high-frequency monitoring data to forecast in-stream turbidity levels in the Catskill Mountains, NY, USA. Results indicate LightGBM models are capable of relatively skillful predictions, which enable robust forecasts for 1–3 days lead times. LightGBM models offer improvements over a simplified linear model across the entire forecast horizon, and more spatially complex models are more resilient to error at shorter lead times (1–3 days). Moreover, interpretation of model features emphasizes high flows as a driver of turbidity in the region. Results suggest that interpretable, flexible, and efficient machine learning algorithms can produce capable water quality forecasts from streamflow forecasts and expand understanding of process dynamics. The use case illustrated here—to our knowledge the first NWM-based water quality forecast—underscores the potential to employ the NWM to expand national water quality forecasting capacity and can overall serve as a guide for similar efforts in basins across the country.

New York

Presence-only modeling using MAXENT: when can we trust the inferences?

1. Recently, interest in species distribution modelling has increased following the development of new methods for the analysis of presence-only data and the deployment of these methods in user-friendly and powerful computer programs. However, reliable inference from these powerful tools requires that several assumptions be met, including the assumptions that observed presences are the consequence of random or representative sampling and that detectability during sampling does not vary with the covariates that determine occurrence probability. 2. Based on our interactions with researchers using these tools, we hypothesized that many presence-only studies were ignoring important assumptions of presence-only modelling. We tested this hypothesis by reviewing 108 articles published between 2008 and 2012 that used the MAXENT algorithm to analyse empirical (i.e. not simulated) data. We chose to focus on these articles because MAXENT has been the most popular algorithm in recent years for analysing presence-only data. 3. Many articles (87%) were based on data that were likely to suffer from sample selection bias; however, methods to control for sample selection bias were rarely used. In addition, many analyses (36%) discarded absence information by analysing presence–absence data in a presence-only framework, and few articles (14%) mentioned detection probability. We conclude that there are many misconceptions concerning the use of presence-only models, including the misunderstanding that MAXENT, and other presence-only methods, relieve users from the constraints of survey design. 4. In the process of our literature review, we became aware of other factors that raised concerns about the validity of study conclusions. In particular, we observed that 83% of articles studies focused exclusively on model output (i.e. maps) without providing readers with any means to critically examine modelled relationships and that MAXENT's logistic output was frequently (54% of articles) and incorrectly interpreted as occurrence probability. 5. We conclude with a series of recommendations foremost that researchers analyse data in a presence–absence framework whenever possible, because fewer assumptions are required and inferences can be made about clearly defined parameters such as occurrence probability.

Methods in Ecology and Evolution

A new framework for analysing automated acoustic species detection data: Occupancy estimation and optimization of recordings post-processing

The development and use of automated species-detection technologies, such as acoustic recorders, for monitoring wildlife are rapidly expanding. Automated classification algorithms provide a cost- and time-effective means to process information-rich data, but often at the cost of additional detection errors. Appropriate methods are necessary to analyse such data while dealing with the different types of detection errors. We developed a hierarchical modelling framework for estimating species occupancy from automated species-detection data. We explore design and optimization of data post-processing procedures to account for detection errors and generate accurate estimates. Our proposed method accounts for both imperfect detection and false positive errors and utilizes information about both occurrence and abundance of detections to improve estimation. Using simulations, we show that our method provides much more accurate estimates than models ignoring the abundance of detections. The same findings are reached when we apply the methods to two real datasets on North American frogs surveyed with acoustic recorders. When false positives occur, estimator accuracy can be improved when a subset of detections produced by the classification algorithm is post-validated by a human observer. We use simulations to investigate the relationship between accuracy and effort spent on post-validation, and found that very accurate occupancy estimates can be obtained with as little as 1% of data being validated. Automated monitoring of wildlife provides opportunity and challenges. Our methods for analysing automated species-detection data help to meet key challenges unique to these data and will prove useful for many wildlife monitoring programs.

Methods in Ecology and Evolution