Geology Reports⌕ Search

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

The effects of topographic surveying technique and data resolution on the detection and interpretation of geomorphic change

Change detection of high resolution topographic data is commonly used in river valleys to quantify reach- and site-scale sediment budgets by estimating the erosion/deposition volume, and to interpret the geomorphic processes driving erosion and deposition. Field survey data are typically collected as point clouds that are often converted to gridded raster datasets and the ultimate choice of grid resolution is left to the user. This choice may have important implications for both the quantification and interpretation of geomorphic change. Here we used concurrent topographic data collected by terrestrial laser scanning (TLS) and structure-from-motion (SfM) photogrammetry to quantify the influence of grid resolution and sampling technique on (a) the sediment budget and (b) the presence and role of geomorphic processes (i.e., alluvial, colluvial, aeolian, and fluvial transport) driving topographic change at four sites along the Colorado River in Grand Canyon, Arizona, USA. We found that while both techniques produced similar estimates for site-scale sediment budgets, the magnitude of detected topographic change was dampened at coarser pixel resolutions. An overall decrease in the areal extent of erosion and deposition were observed, respectively, when coarsening pixel size from 5 cm to 1 m among all sites. Coarser resolution data tended to affect interpretation of landscape change along the margins of river valleys. For example, when changing from 5 cm to 1 m pixel resolution, the inferred contribution of aeolian changes to total site-scale geomorphic change increased in area by 7.9%, whereas the inferred contribution of alluvial and colluvial processes decreased in area by 97.9% and 88.2%, respectively. More generally, we found that coarsening pixel sizes disproportionately attributed geomorphic change to one or more of the most common processes operating at a site. We also found that coarsening pixel resolution amplified the net sediment imbalance at the site scale, driving the imbalance at erosional sites further into erosion and vice versa for depositional sites. Our results have implications both for point cloud data collection and for raster dataset processing. We argue that selecting the finest obtainable resolution is not always warranted to accurately quantify and interpret geomorphic change, because remote sensing technique, topographic data resolution, and analysis procedure can be optimized to capture the spatial scale of those processes driving landscape change. However, in landscapes at or near sediment equilibrium (i.e., equal amounts of erosion and deposition), the finest obtainable topographic data resolution is warranted to avoid amplifying sediment imbalance and erroneously inferring that sites are trending toward erosion or deposition.

Arizona↗

Optimizing selection of training and auxiliary data for operational land cover classification for the LCMAP initiative

The U.S. Geological Survey’s Land Change Monitoring, Assessment, and Projection (LCMAP) initiative is a new end-to-end capability to continuously track and characterize changes in land cover, use, and condition to better support research and applications relevant to resource management and environmental change. Among the LCMAP product suite are annual land cover maps that will be available to the public. This paper describes an approach to optimize the selection of training and auxiliary data for deriving the thematic land cover maps based on all available clear observations from Landsats 4–8. Training data were selected from map products of the U.S. Geological Survey’s Land Cover Trends project. The Random Forest classifier was applied for different classification scenarios based on the Continuous Change Detection and Classification (CCDC) algorithm. We found that extracting training data proportionally to the occurrence of land cover classes was superior to an equal distribution of training data per class, and suggest using a total of 20,000 training pixels to classify an area about the size of a Landsat scene. The problem of unbalanced training data was alleviated by extracting a minimum of 600 training pixels and a maximum of 8000 training pixels per class. We additionally explored removing outliers contained within the training data based on their spectral and spatial criteria, but observed no significant improvement in classification results. We also tested the importance of different types of auxiliary data that were available for the conterminous United States, including: (a) five variables used by the National Land Cover Database, (b) three variables from the cloud screening ‘‘Function of mask” (Fmask) statistics, and (c) two variables from the change detection results of CCDC. We found that auxiliary variables such as a Digital Elevation Model and its derivatives (aspect, position index, and slope), potential wetland index, water probability, snow probability, and cloud probability improved the accuracy of land cover classification. Compared to the original strategy of the CCDC algorithm (500 pixels per class), the use of the optimal strategy improved the classification accuracies substantially (15-percentage point increase in overall accuracy and 4-percentage point increase in minimum accuracy).

ISPRS Journal of Photogrammetry and Remote Sensing↗

Accuracy assessment of NLCD 2011 impervious cover data for the Chesapeake Bay region, USA

The National Land Cover Database (NLCD) contains three eras (2001, 2006, 2011) of percentage urban impervious cover (%IC) at the native pixel size (30 m-x-30 m) of the Landsat Thematic Mapper satellite. These data are potentially valuable to environmental managers and stakeholders because of the utility of %IC as an indicator of watershed and aquatic condition, but lack an accuracy assessment because of the absence of suitable reference data. Recently developed 1 m2 land cover data for the Chesapeake Bay region makes it possible to assess NLCD %IC accuracy for a 262,000 km2 region based on a census rather than a sample of reference data. We report agreement between the two %IC datasets for watersheds and the riparian zones within watersheds and four additional square units. The areas of the six assessment units were 40 ha cell, 433 ha (riparian unit average), 2,756 ha cell, 5,626 ha cell, 8,569 ha (watershed unit average) and 22,500 ha cell. Mean Absolute Deviation (MAD) was ≤ 1.6% for each of the six assessment units and Mean Deviation (MD) was only slightly less, indicating NLCD reliably reproduced %IC from the 1 m2 data with a small (≤ 1.6%) and consistent tendency for underestimation. Results were sensitive to assessment unit choice. The results for the four largest assessment units had very similar regression parameters, R2 values, and patterns of bias. Results for the riparian assessment were different from those for the watershed unit and the other three larger units. MAD was about 50% less for the riparian zones than it was for the watersheds, the direction of bias was less consistent, and NLCD %IC was uniformly higher than 1 m2 %IC in urbanized riparian zones. For the smallest unit, bias patterns were more similar to the riparian unit and regression results were more similar to the four larger units. MAD and MD were also sensitive to the amount of urbanization, increasing as NLCD %IC increased. The low overall bias and positive relationship between bias and level of urbanization suggest that the benefits of obtaining 1 m2 IC data outside of urban areas may not outweigh the costs of obtaining such data.

Delaware, Maryland, New York, Pennsylvania, Virgin↗

Slip distribution and rupture history of the August 11, 2012, double earthquakes in Ahar – Varzaghan, Iran, using joint inversion of teleseismic broadband and local strong motion data

We use combined teleseismic and strong motion data sets to investigate finite-fault slip models for a double of earthquakes that occurred on August 11, 2012, in northwestern Iran near the cities of Ahar and Varzaghan. The data include teleseismic P-waveforms retrieved from broadband seismic stations located between 30°–94° from the earthquakes and local strong motion data recorded by the Iran Strong Motion Network, installed and operated by the Building and Housing Research Centre. We first invert teleseismic P-waveforms and local strong motion data separately. For the first event (12:23 UTC), the teleseismic broadband inversion yields a somewhat deeper and simpler distribution of slip than the local strong motion inversion. The strong motion inversion results in a more complex distribution because of higher frequency content but can also be influenced by complexities in the propagation path. For the second event (12:34 UTC), the slip distribution from strong motion data is more similar to the teleseismic result and shows a simple slip area with a small relative movement to the west. To resolve the differences between the results of these two data sets and obtain a better constrained slip model, we perform a joint inversion of teleseismic broadband and local strong motion data. The joint inversion for the first event shows two asperities with a maximum slip of 3.9 m up- dip from the hypocenter and extending to the west between depths of 1 and 5 km. A second narrower high-slip area is seen just above the hypocenter from 6 to 10 km depth. The total moment for this earthquake is calculated to be M o = 3.8 × 10 25 dyn-cm (3.8 × 10 18 N.m) (M w 6.4). For the second event, the results of the joint inversion show a simple slip distribution that is mainly confined in a single patch around the hypocenter with a depth range from about 10 to 13 km and maximum slip of 1.9 m. We compute a total seismic moment of M o = 1.6 × 10 25 dyn-cm (1.6 × 10 18 N.m) (M w 6.1) for the second event. The largest stress drops for the first event occur above the hypocenter with an average stress drop over the rupture area of 120 bar (12 Mpa). For the second event, the maximum stress drop occurs at the reported focal depth with an average stress drop over the rupture area of 80 bar (8 Mpa).

East Anatolian Fault↗

Interoperability in planetary research for geospatial data analysis

For more than a decade there has been a push in the planetary science community to support interoperable methods for accessing and working with geospatial data. Common geospatial data products for planetary research include image mosaics, digital elevation or terrain models, geologic maps, geographic location databases (e.g., craters, volcanoes) or any data that can be tied to the surface of a planetary body (including moons, comets or asteroids). Several U.S. and international cartographic research institutions have converged on mapping standards that embrace standardized geospatial image formats, geologic mapping conventions, U.S. Federal Geographic Data Committee (FGDC) cartographic and metadata standards, and notably on-line mapping services as defined by the Open Geospatial Consortium (OGC). The latter includes defined standards such as the OGC Web Mapping Services (simple image maps), Web Map Tile Services (cached image tiles), Web Feature Services (feature streaming), Web Coverage Services (rich scientific data streaming), and Catalog Services for the Web (data searching and discoverability). While these standards were developed for application to Earth-based data, they can be just as valuable for planetary domain. Another initiative, called VESPA (Virtual European Solar and Planetary Access), will marry several of the above geoscience standards and astronomy-based standards as defined by International Virtual Observatory Alliance (IVOA). This work outlines the current state of interoperability initiatives in use or in the process of being researched within the planetary geospatial community.

Planetary and Space Science↗

Mapping burned areas using dense time-series of Landsat data

Complete and accurate burned area data are needed to document patterns of fires, to quantify relationships between the patterns and drivers of fire occurrence, and to assess the impacts of fires on human and natural systems. Unfortunately, in many areas existing fire occurrence datasets are known to be incomplete. Consequently, the need to systematically collect burned area information has been recognized by the United Nations Framework Convention on Climate Change and the Intergovernmental Panel on Climate Change, which have both called for the production of essential climate variables (ECVs), including information about burned area. In this paper, we present an algorithm that identifies burned areas in dense time-series of Landsat data to produce the Landsat Burned Area Essential Climate Variable (BAECV) products. The algorithm uses gradient boosted regression models to generate burn probability surfaces using band values and spectral indices from individual Landsat scenes, lagged reference conditions, and change metrics between the scene and reference predictors. Burn classifications are generated from the burn probability surfaces using pixel-level thresholding in combination with a region growing process. The algorithm can be applied anywhere Landsat and training data are available. For this study, BAECV products were generated for the conterminous United States from 1984 through 2015. These products consist of pixel-level burn probabilities for each Landsat scene, in addition to, annual composites including: the maximum burn probability and a burn classification. We compared the BAECV burn classification products to the existing Global Fire Emissions Database (GFED; 1997–2015) and Monitoring Trends in Burn Severity (MTBS; 1984–2013) data. We found that the BAECV products mapped 36% more burned area than the GFED and 116% more burned area than MTBS. Differences between the BAECV products and the GFED were especially high in the West and East where the BAECV products mapped 32% and 88% more burned area, respectively. However, the BAECV products found less burned area than the GFED in regions with frequent agricultural fires. Compared to the MTBS data, the BAECV products identified 31% more burned area in the West, 312% more in the Great Plains, and 233% more in the East. Most pixels in the MTBS data were detected by the BAECV, regardless of burn severity. The BAECV products document patterns of fire similar to those in the GFED but also showed patterns of fire that are not well characterized by the existing MTBS data. We anticipate the BAECV products will be useful to studies that seek to understand past patterns of fire occurrence, the drivers that created them, and the impacts fires have on natural and human systems.

continental United States↗

Recovering individual-level spatial inference from aggregated binary data

Binary regression models are commonly used in disciplines such as epidemiology and ecology to determine how spatial covariates influence individuals. In many studies, binary data are shared in a spatially aggregated form to protect privacy. For example, rather than reporting the location and result for each individual that was tested for a disease, researchers may report that a disease was detected or not detected within geopolitical units. Often, the spatial aggregation process obscures the values of response variables, spatial covariates, and locations of each individual, which makes recovering individual-level inference difficult. We show that applying a series of transformations, including a change of support, to a bivariate point process model allows researchers to recover individual-level inference for spatial covariates from spatially aggregated binary data. The series of transformations preserves the convenient interpretation of desirable binary regression models that are commonly applied to individual-level data. Using a simulation experiment, we compare the performance of our proposed method under varying types of spatial aggregation against the performance of standard approaches using the original individual-level data. We illustrate our method by modeling individual-level probability of infection using a data set that has been aggregated to protect an at-risk and endangered species of bats. Our simulation experiment and data illustration demonstrate the utility of the proposed method when access to original non-aggregated data is impractical or prohibited.

Northeast and Midwest United States↗

Standardized data quality acceptance criteria for a rapid Escherichia coli qPCR method (Draft Method C) for water quality monitoring at recreational beaches

There is growing interest in the application of rapid quantitative polymerase chain reaction (qPCR) and other PCR-based methods for recreational water quality monitoring and management programs. This interest has strengthened given the publication of U.S. Environmental Protection Agency (EPA)-validated qPCR methods for enterococci fecal indicator bacteria (FIB) and has extended to similar methods for Escherichia coli ( E. coli ) FIB. Implementation of qPCR-based methods in monitoring programs can be facilitated by confidence in the quality of the data produced by these methods. Data quality can be determined through the establishment of a series of specifications that should reflect good laboratory practice. Ideally, these specifications will also account for the typical variability of data coming from multiple users of the method. This study developed proposed standardized data quality acceptance criteria that were established for important calibration model parameters and/or controls from a new qPCR method for E. coli (EPA Draft Method C) based upon data that was generated by 21 laboratories. Each laboratory followed a standardized protocol utilizing the same prescribed reagents and reference and control materials. After removal of outliers, statistical modeling based on a hierarchical Bayesian method was used to establish metrics for assay standard curve slope, intercept and lower limit of quantification that included between-laboratory, replicate testing within laboratory, and random error variability. A nested analysis of variance (ANOVA) was used to establish metrics for calibrator/positive control, negative control, and replicate sample analysis data. These data acceptance criteria should help those who may evaluate the technical quality of future findings from the method, as well as those who might use the method in the future. Furthermore, these benchmarks and the approaches described for determining them may be helpful to method users seeking to establish comparable laboratory-specific criteria if changes in the reference and/or control materials must be made.

Water Research↗

Characteristic length scale of input data in distributed models: implications for modeling grain size

The appropriate spatial scale for a distributed energy balance model was investigated by: (a) determining the scale of variability associated with the remotely sensed and GIS-generated model input data; and (b) examining the effects of input data spatial aggregation on model response. The semi-variogram and the characteristic length calculated from the spatial autocorrelation were used to determine the scale of variability of the remotely sensed and GIS-generated model input data. The data were collected from two hillsides at Upper Sheep Creek, a sub-basin of the Reynolds Creek Experimental Watershed, in southwest Idaho. The data were analyzed in terms of the semivariance and the integral of the autocorrelation. The minimum characteristic length associated with the variability of the data used in the analysis was 15 m. Simulated and observed radiometric surface temperature fields at different spatial resolutions were compared. The correlation between agreement simulated and observed fields sharply declined after a 10×10 m 2 modeling grid size. A modeling grid size of about 10×10 m 2 was deemed to be the best compromise to achieve: (a) reduction of computation time and the size of the support data; and (b) a reproduction of the observed radiometric surface temperature.

Idaho↗

Classification methods for monitoring Arctic sea ice using OKEAN passive/active two-channel microwave data

This paper presents methods for classifying Arctic sea ice using both passive and active (2-channel) microwave imagery acquired by the Russian OKEAN 01 polar-orbiting satellite series. Methods and results are compared to sea ice classifications derived from nearly coincident Special Sensor Microwave Imager (SSM/I) and Advanced Very High Resolution Radiometer (AVHRR) image data of the Barents, Kara, and Laptev Seas. The Russian OKEAN 01 satellite data were collected over weekly intervals during October 1995 through December 1997. Methods are presented for calibrating, georeferencing and classifying the raw active radar and passive microwave OKEAN 01 data, and for correcting the OKEAN 01 microwave radiometer calibration wedge based on concurrent 37 GHz horizontal polarization SSM/I brightness temperature data. Sea ice type and ice concentration algorithms utilized OKEAN's two-channel radar and passive microwave data in a linear mixture model based on the measured values of brightness temperature and radar backscatter, together with a priori knowledge about the scattering parameters and natural emissivities of basic sea ice types. OKEAN 01 data and algorithms tended to classify lower concentrations of young or first-year sea ice when concentrations were less than 60%, and to produce higher concentrations of multi-year sea ice when concentrations were greater than 40%, when compared to estimates produced from SSM/I data. Overall, total sea ice concentration maps derived independently from OKEAN 01, SSM/I, and AVHRR satellite imagery were all highly correlated, with uniform biases, and mean differences in total ice concentration of less than four percent (sd<15%).

Remote Sensing of Environment↗

Seasonal comparisons of sea ice concentration estimates derived from SSM/I, OKEAN, and RADARSAT data

The Special Sensor Microwave Imager (SSM/I) microwave satellite radiometer and its predecessor SMMR are primary sources of information for global sea ice and climate studies. However, comparisons of SSM/I, Landsat, AVHRR, and ERS-1 synthetic aperture radar (SAR) have shown substantial seasonal and regional differences in their estimates of sea ice concentration. To evaluate these differences, we compared SSM/I estimates of sea ice coverage derived with the NASA Team and Bootstrap algorithms to estimates made using RADARSAT, and OKEAN-01 satellite sensor data. The study area included the Barents Sea, Kara Sea, Laptev Sea, and adjacent parts of the Arctic Ocean, during October 1995 through October 1999. Ice concentration estimates from spatially and temporally near-coincident imagery were calculated using independent algorithms for each sensor type. The OKEAN algorithm implemented the satellite's two-channel active (radar) and passive microwave data in a linear mixture model based on the measured values of brightness temperature and radar backscatter. The RADARSAT algorithm utilized a segmentation approach of the measured radar backscatter, and the SSM/I ice concentrations were derived at National Snow and Ice Data Center (NSIDC) using the NASA Team and Bootstrap algorithms. Seasonal and monthly differences between SSM/I, OKEAN, and RADARSAT ice concentrations were calculated and compared. Overall, total sea ice concentration estimates derived independently from near-coincident RADARSAT, OKEAN-01, and SSM/I satellite imagery demonstrated mean differences of less than 5.5% (S.D.<9.5%) during the winter period. Differences between the SSM/I NASA Team and the SSM/I Bootstrap concentrations were no more than 3.1% (S.D.<5.4%) during this period. RADARSAT and OKEAN-01 data both yielded higher total ice concentrations than the NASA Team and the Bootstrap algorithms. The Bootstrap algorithm yielded higher total ice concentrations than the NASA Team algorithm. Total ice concentrations derived from OKEAN-01 and SSM/I satellite imagery were highly correlated during winter, spring, and fall, with mean differences of less than 8.1% (S.D.<15%) for the NASA Team algorithm, and less than 2.8% (S.D.<13.8%) for the Bootstrap algorithm. Respective differences between SSM/I NASA Team and SSM/I Bootstrap total concentrations were less than 5.3% (S.D.<6.9%). Monthly mean differences between SSM/I and OKEAN differed annually by less than 6%, with smaller differences primarily in winter. The NASA Team and Bootstrap algorithms underestimated the total sea ice concentrations relative to the RADARSAT ScanSAR no more than 3.0% (S.D.<9%) and 1.2% (S.D.<7.5%) during cold months, and no more than 12% and 7% during summer, respectively. ScanSAR tended to estimate higher ice concentrations for ice concentrations greater than 50%, when compared to SSM/I during all months. ScanSAR underestimated total sea ice concentration by 2% compared to the OKEAN-01 algorithm during cold months, and gave an overestimation by 2% during spring and summer months. Total NASA Team and Bootstrap sea ice concentration estimates derived from coincident SSM/I and OKEAN-01 data demonstrated mean differences of no more than 5.3% (S.D.<7%), 3.1% (S.D.<5.5%), 2.0% (S.D.<5.5%), and 7.3% (S.D.<10%) for fall, winter, spring, and summer periods, respectively. Large disagreements were observed between the OKEAN and NASA Team results in spring and summer for estimates of the first-year (FY) and multiyear (MY) age classes. The OKEAN-01 algorithm and data tended to estimate, on average, lower concentrations of young or FY ice and higher concentrations of total and MY ice for all months and seasons. Our results contribute to the growing body of documentation about the levels of disparity obtained when seasonal sea ice concentrations are estimated using various types of satellite data and algorithms.

Remote Sensing of Environment↗

Paleomagnetic data bearing on style of Miocene deformation in the Lake Mead area, Southern Nevada

Paleomagnetic and structural data from intermediate to mafic composition lava flows and related dikes in all major blocks of the late Miocene Hamblin-Cleopatra Volcano, which was structurally dismembered during the development of the Lake Mead Fault System (LMFS), provide limits on the magnitude and sense of tilting and vertical axis rotation of crust during extension of this part of the Basin and Range province. Sinistral separation along the fault system dissected the volcano into three major blocks. The eastern, Cleopatra Lobe of the volcano is structurally the most intact section of the volcano. Normal and reverse polarity data from paleomagnetic sites collected along traverses in the Cleopatra Lobe yield an in situ grand mean of Declination (D) = 339??, Inclination (I) = +54??, ??95 = 3.1??, k = 27.2, N = 81 sites. The rocks of the central core of the volcano yield an in situ grand mean of D = 3??, I = + 59??, ??95 = 6.8??, k = 42.5, N = 11 sites (six normal, five reverse polarity). Sites collected within the western Hamblin Lobe of the volcano are exclusively of reverse polarity and yield an overall in situ mean of D = 168??, I = -58??, ??95 = 6.5??. k = 28.9, N = 18 sites. Interpretation of the paleomagnetic data in the context of the structural history of the volcano and surrounding area, considers the possibility of two different types of structural corrections. A stratigraphic tilt correction involves restoring flows to the horizontal using the present strike. This correction assumes no initial, possibly radial, dip of flows of the volcano and is considered invalid. A structural tilt correction to the data assumes that dikes of the radiating swarm associated with the volcano were originally vertical and results in block mean directions of D = 9??, I = +53??, ??95 = 3.1??, k = 27.2, and D = 58??, I = + 78??, ??95 = 6.8, k = 42.5, for the Cleopatra Lobe and the central intrusive core, respectively. The data from the Cleopatra Lobe are slightly discordant, in a clockwise sense, from expected middle- to late-Miocene field directions. The data from the volcano are not consistent with a proposed structural model of uniform, moderate magnitude, statistically significant, counter-clockwise vertical axis rotation of fault-bounded blocks during overall sinsitral displacement along the LMFS. We also analyzed dikes of the northernmost part of the Miocene Wilson Ridge hypabyssal igneous complex, strata of the Triassic Chinle Formation, and basalt flows of the Miocene West End Wash/Callville Mesa volcanic centers. Dikes in the Wilson Ridge pluton and the Triassic strata yield magnetizations with directions suggestive of statistically significant, clockwise, vertical-axis rotations consistent with local, large-magnitude shear of crustal fragments near some of the faults of the LMFS. Late Cenozoic deformation of the Hamblin-Cleopatra volcano area appears to have been non-uniform in scale and magnitude and no single structural model, involving strictly strike-slip faulting, can account for the observed paleomagnetic data. ?? 2001 Elsevier Science Ltd. All rights reserved.

Journal of Structural Geology↗

Analysis of environmental data with censored observations

The potential threats to humans and to terrestrial and aquatic ecosystems from environmental contamination could depend on the sum of the concentrations of different chemicals. However, direct summation of environmental data is not generally feasible because it is common for some chemical concentrations to be recorded as being below the analytical reporting limit. This creates special problems in the analysis of the data. A new model selection procedure, named forward censored regression, is introduced for selecting an appropriate model for environmental data with censored observations. The procedure is demonstrated using concentrations of atrazine (2-chloro-4-ethylamino-6-isopropylamino- s -triazine), deethylatrazine (DEA, 2-amino-4-chloro-6-isopropylamino- s -triazine), and deisopropylatrazine (DIA, 2-amino-4-chloro-6-ethylamino- s -triazine) in groundwater in the midwestern United States by using the data derived from a previous study conducted by the U.S. Geological Survey. More than 80% of the observations for each compound for this study were left censored at 0.05 &mu;g/L. The values for censored observations of atrazine, DEA, and DIA are imputed with the selected models. The summation of atrazine residue (atrazine + DEA + DIA) can then be calculated using the combination of observed and imputed values to generate a pseudo-complete data set. The all-subsets regression procedure is applied to the pseudo-complete data to select the final model for atrazine residue. The methodology presented can be used to analyze similar cases of environmental contamination involving censored data.

Environmental Science & Technology↗

Some simple guides to finding useful information in exploration geochemical data

Most regional geochemistry data reflect processes that can produce superfluous bits of noise and, perhaps, information about the mineralization process of interest. There are two end-member approaches to finding patterns in geochemical data—unsupervised learning and supervised learning. In unsupervised learning, data are processed and the geochemist is given the task of interpreting and identifying possible sources of any patterns. In supervised learning, data from known subgroups such as rock type, mineralized and nonmineralized, and types of mineralization are used to train the system which then is given unknown samples to classify into these subgroups. To locate patterns of interest, it is helpful to transform the data and to remove unwanted masking patterns. With trace elements use of a logarithmic transformation is recommended. In many situations, missing censored data can be estimated using multiple regression of other uncensored variables on the variable with censored values. In unsupervised learning, transformed values can be standardized, or normalized, to a Z-score by subtracting the subset's mean and dividing by its standard deviation. Subsets include any source of differences that might be related to processes unrelated to the target sought such as different laboratories, regional alteration, analytical procedures, or rock types. Normalization removes effects of different means and measurement scales as well as facilitates comparison of spatial patterns of elements. These adjustments remove effects of different subgroups and hopefully leave on the map the simple and uncluttered pattern(s) related to the mineralization only. Supervised learning methods, such as discriminant analysis and neural networks, offer the promise of consistent and, in certain situations, unbiased estimates of where mineralization might exist. These methods critically rely on being trained with data that encompasses all populations fairly and that can possibly fall into only the identified populations.

Natural Resources Research↗

Singularity and Nonnormality in the Classification of Compositional Data

Geologists may want to classify compositional data and express the classification as a map. Regionalized classification is a tool that can be used for this purpose, but it incorporates discriminant analysis, which requires the computation and inversion of a covariance matrix. Covariance matrices of compositional data always will be singular (noninvertible) because of the unit-sum constraint. Fortunately, discriminant analyses can be calculated using a pseudo-inverse of the singular covariance matrix; this is done automatically by some statistical packages such as SAS. Granulometric data from the Darss Sill region of the Baltic Sea is used to explore how the pseudo-inversion procedure influences discriminant analysis results, comparing the algorithm used by SAS to the more conventional Moore-Penrose algorithm. Logratio transforms have been recommended to overcome problems associated with analysis of compositional data, including singularity. A regionalized classification of the Darss Sill data after logratio transformation is different only slightly from one based on raw granulometric data, suggesting that closure problems do not influence severely regionalized classification of compositional data.

Mathematical Geology↗

Comparisons of two moments‐based estimators that utilize historical and paleoflood data for the log Pearson type III distribution

The expected moments algorithm (EMA) [ Cohn et al. , 1997 ] and the Bulletin 17B [ Interagency Committee on Water Data , 1982 ] historical weighting procedure (B17H) for the log Pearson type III distribution are compared by Monte Carlo computer simulation for cases in which historical and/or paleoflood data are available. The relative performance of the estimators was explored for three cases: fixed‐threshold exceedances, a fixed number of large floods, and floods generated from a different parent distribution. EMA can effectively incorporate four types of historical and paleoflood data: floods where the discharge is explicitly known, unknown discharges below a single threshold, floods with unknown discharge that exceed some level, and floods with discharges described in a range. The B17H estimator can utilize only the first two types of historical information. Including historical/paleoflood data in the simulation experiments significantly improved the quantile estimates in terms of mean square error and bias relative to using gage data alone. EMA performed significantly better than B17H in nearly all cases considered. B17H performed as well as EMA for estimating X 100 in some limited fixed‐threshold exceedance cases. EMA performed comparatively much better in other fixed‐threshold situations, for the single large flood case, and in cases when estimating extreme floods equal to or greater than X 500 . B17H did not fully utilize historical information when the historical period exceeded 200 years. Robustness studies using GEV‐simulated data confirmed that EMA performed better than B17H. Overall, EMA is preferred to B17H when historical and paleoflood data are available for flood frequency analysis.

Water Resources Research↗

Monitoring eruptive activity at Mount St. Helens with TIR image data

Thermal infrared (TIR) data from the MASTER airborne imaging spectrometer were acquired over Mount St. Helens in Sept and Oct, 2004, before and after the onset of recent eruptive activity. Pre‐eruption data showed no measurable increase in surface temperatures before the first phreatic eruption on Oct 1. MASTER data acquired during the initial eruptive episode on Oct 14 showed maximum temperatures of ∼330°C and TIR data acquired concurrently from a Forward Looking Infrared (FLIR) camera showed maximum temperatures ∼675°C, in narrow (∼1‐m) fractures of molten rock on a new resurgent dome. MASTER and FLIR thermal flux calculations indicated a radiative cooling rate of ∼714 J/m 2 /s over the new dome, corresponding to a radiant power of ∼24 MW. MASTER data indicated the new dome was dacitic in composition, and digital elevation data derived from LIDAR acquired concurrently with MASTER showed that the dome growth correlated with the areas of elevated temperatures. Low SO 2 concentrations in the plume combined with sub‐optimal viewing conditions prohibited quantitative measurement of plume SO 2 . The results demonstrate that airborne TIR data can provide information on the temperature of both the surface and plume and the composition of new lava during eruptive episodes. Given sufficient resources, the airborne instrumentation could be deployed rapidly to a newly‐awakening volcano and provide a means for remote volcano monitoring.

Washington↗

Did they feel it? Legacy maroseismic data illuminates an engimatic 20th century earthquake

The challenges and the importance of preserving legacy instrumental records of earthquakes are now well-recognized (e.g., Richards & Hellweg, 2020, https://doi.org/10.1785/0220200053 ). Seismologists may not be aware of parallel challenges and opportunities with legacy macroseismic data for earthquakes in the United States. For much of the 20th century, macroseismic data were collected by a series of U.S. government agencies using a standard questionnaire distributed on postcards. Published summaries of postcards provide macroseismic data akin to modern Did You Feel It? questionnaire responses. In this paper we focus on the M 6.5 Fickle Hill, California earthquake, on 21 December 1954 (Hellweg et al., 2025) as a proof-of-concept, illustrating the potential of what we dub Did They Feel It? (DTFI) data to improve our understanding of significant 20th century U.S. earthquakes for which instrumental data are sparse. Legacy macroseismic data interpreted following modern conventions can potentially constrain traditional ShakeMaps at a level of detail and accuracy that in some respects rival maps for modern earthquakes. The updated ShakeMap for the 1954 Fickle Hill earthquake, also drawing from recently published media and first-person accounts, supports the location, depth, and stress drop value estimated from available instrumental data (Hellweg et al., 2025).

California↗