Geology Reports⌕ Search

SEARCH · Geology Reports

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Effects of model complexity and priors on estimation using sequential importance sampling/resampling for species conservation

We examined the effects of complexity and priors on the accuracy of models used to estimate ecological and observational processes, and to make predictions regarding population size and structure. State-space models are useful for estimating complex, unobservable population processes and making predictions about future populations based on limited data. To better understand the utility of state space models in evaluating population dynamics, we used them in a Bayesian framework and compared the accuracy of models with differing complexity, with and without informative priors using sequential importance sampling/resampling (SISR). Count data were simulated for 25 years using known parameters and observation process for each model. We used kernel smoothing to reduce the effect of particle depletion, which is common when estimating both states and parameters with SISR. Models using informative priors estimated parameter values and population size with greater accuracy than their non-informative counterparts. While the estimates of population size and trend did not suffer greatly in models using non-informative priors, the algorithm was unable to accurately estimate demographic parameters. This model framework provides reasonable estimates of population size when little to no information is available; however, when information on some vital rates is available, SISR can be used to obtain more precise estimates of population size and process. Incorporating model complexity such as that required by structured populations with stage-specific vital rates affects precision and accuracy when estimating latent population variables and predicting population dynamics. These results are important to consider when designing monitoring programs and conservation efforts requiring management of specific population segments.

Ecological Modelling↗

A three-dimensional Lagrangian particle tracking model for predicting transport of eggs of rheophilic-spawning carps in turbulent rivers

Grass carp, bighead carp, and silver carp spawn in flowing water. Their eggs, and then larvae, develop while drifting. Hydraulic conditions and water temperature control spawning locations, egg survival, and the downstream distance traveled before the hatched larvae can swim for low velocity nursery habitats. Existing egg drift models simulate the fluvial transport of carp eggs but have limitations in capturing the effect of localized turbulence on egg transport due to inadequate dimensions of hydrodynamics and/or empirical parameterization of river dispersion. We present a three-dimensional Lagrangian particle tracking model that uses fully resolved river hydrodynamics and a continuous random walk algorithm driven by local turbulent kinetic energy and its dissipation rate. We incorporate a new set of equations to compute evolving egg characteristics with fully resolved 3-D hydrodynamics. To demonstrate the performance of the model, we conducted a case study in an eight-kilometer reach of the Missouri River at the discharge of approximately 25% daily flow exceedance. Three-dimensional river hydrodynamics was modeled, calibrated, and evaluated with measurement data. Egg drift was modeled and compared using fully three-dimensional, depth-averaged two-dimensional, and zone-averaged one-dimensional hydrodynamics. The comparison shows a generally good agreement among models of downstream egg transport due to advection but a different dispersion pattern of eggs in the river, as a result of turbulent diffusion and shear induced dispersion.

Ecological Modelling↗

Machine learning for ecosystem services

Recent developments in machine learning have expanded data-driven modelling (DDM) capabilities, allowing artificial intelligence to infer the behaviour of a system by computing and exploiting correlations between observed variables within it. Machine learning algorithms may enable the use of increasingly available ‘big data’ and assist applying ecosystem service models across scales, analysing and predicting the flows of these services to disaggregated beneficiaries. We use the Weka and ARIES software to produce two examples of DDM: firewood use in South Africa and biodiversity value in Sicily, respectively. Our South African example demonstrates that DDM (64–91% accuracy) can identify the areas where firewood use is within the top quartile with comparable accuracy as conventional modelling techniques (54–77% accuracy). The Sicilian example highlights how DDM can be made more accessible to decision makers, who show both capacity and willingness to engage with uncertainty information. Uncertainty estimates, produced as part of the DDM process, allow decision makers to determine what level of uncertainty is acceptable to them and to use their own expertise for potentially contentious decisions. We conclude that DDM has a clear role to play when modelling ecosystem services, helping produce interdisciplinary models and holistic solutions to complex socio-ecological issues.

Ecosystem Services↗

Assessing the impacts of climate and land use and land cover change on the freshwater availability in the Brahmaputra River basin

Study Region: Brahmaputra River basin in South Asia. Study Focus: The Soil and Water Assessment Tool was used to evaluate sensitivities and patterns in freshwater availability due to projected climate and land use changes in the Brahmaputra basin. The daily observed discharge at Bahadurabad station in Bangladesh was used to calibrate and validate the model and analyze uncertainties with a sequential uncertainty fitting algorithm. The sensitivities and impacts of projected climate and land use changes on basin hydrological components were simulated for the A1B and A2 scenarios and analyzed relative to a baseline scenario of 1988–2004. New hydrological insights for the region: Basin average annual ET was found to be sensitive to changes in CO 2 concentration and temperature, while total water yield, streamflow, and groundwater recharge were sensitive to changes in precipitation. The basin hydrological components were predicted to increase with seasonal variability in response to climate and land use change scenarios. Strong increasing trends were predicted for total water yield, streamflow, and groundwater recharge, indicating exacerbation of flooding potential during August–October, but strong decreasing trends were predicted, indicating exacerbation of drought potential during May–July of the 21st century. The model has potential to facilitate strategic decision making through scenario generation integrating climate change adaptation and hazard mitigation policies to ensure optimized allocation of water resources under a variable and changing climate.

Brahmaputra River basin↗

Improving crop-specific groundwater use estimation in the Mississippi Alluvial Plain: Implications for integrated remote sensing and machine learning approaches in data-scarce regions

Study region The Mississippi Alluvial Plain (MAP) in the United States (US). Study focus Understanding local-scale groundwater use, a critical component of the water budget, is necessary for implementing sustainable water management practices. The MAP is one of the most productive agricultural regions in the US and extracts more than 11 km 3 /year for irrigation activities. Consequently, groundwater-level declines in the MAP region pose a substantial challenge to water sustainability, and hence, we need reliable groundwater pumping monitoring solutions to manage this resource appropriately. New hydrological insights for the region We incorporate remote sensing datasets and machine learning to improve an existing lookup table-based model of groundwater use previously developed by the U.S. Geological Survey (USGS). Here, we employ Distributed Random Forests, an ensemble machine learning algorithm to predict annual and monthly groundwater use (2014–2020) throughout this region at 1-km resolution, using pumping data from existing flowmeters in the Mississippi Delta. Our model compares favorably with the existing USGS model, with higher R 2 (0.51 compared to 0.42 in the previous model), and lower root mean square error (RMSE) and mean absolute error (MAE)— 0.14 m and 0.09 m, respectively in our model, compared to 0.15 m and 0.1 m in the previous model. Therefore, this work advances our ability to predict groundwater use in regions with scarce or limited in-situ groundwater withdrawal data availability.

Journal of Hydrology Regional Studies↗

Establishment of baseline cytology metrics in nestling American kestrels (Falco sparverius): Immunomodulatory effects of the flame retardant isopropylated triarylphosphate isomers

Avian populations must mount effective immune responses upon exposure to environmental stressors such as avian influenza and xenobiotics. Although multiple immune assays have been tested and applied to various avian species, antibody-mediated immune responses in non-model avian species are not commonly reported due to the lack of commercially available species-specific antibodies. The objectives of the present study were to advance methods for studying wild bird immune responses and to apply these to the evaluation of cytological responses after exposure of American kestrels, Falco sparverius, to a commercial flame retardant mixture containing isopropylated triarylphosphate isomers (ITP). Hatchlings were gavaged daily with safflower oil or 1.5 ug/g bw/day of ITP suspended in safflower oil, then bled on days 9, 17, and 21. The ITP treatment group ( n = 18) and a subset of controls (Poly I:C treatment group; n = 10) were injected on days 9 and 15 with a synthetic analog of viral double-stranded RNA, polyinosinic:polycytidylic acid (Poly I:C), a toll-like receptor ligand and synthetic viral mimic, and responses compared to a sham injected control group (n = 8). The hypotheses tested whether kestrels showed immunological differences among treatment groups, genetic sex, and/or white blood cell (WBC) subpopulation type over time. A flow cytometry (FCM) gating strategy categorized heterophils (H), lymphocytes (L), and monocytes (M) and their proportions, and measured relative fluorescence in response to anti-chicken CD4 binding. Fluorescent cell surfaces and some granular/vacuolar inclusions were visualized by epifluorescence microscopy. A fourth subpopulation with higher levels of granularity than M but less than H became increasingly apparent with time and was gated along with the H subpopulation; its frequency of occurrence was lowest in the ITP group ( P = 0.0023). The percentages of cells differed among treatment groups, days, and sexes ( P = 0.0001). For both sexes, percentages of H and L were higher than M in control and Poly I:C. In the ITP group, L percentages were higher than H and M ( P = 0.0457), and H and L were higher than M on days 9 and 21 ( P = 0.0001). The ratios of H:L and H:WBC, indicators of robust immunity, were also higher on days 9 and 21 than on 17 ( P = 0.0079). For each sex, the highest levels of activity measured by FCM geometric means (GEO) of fluorescence (indicative of antibody binding) were observed on day 9 ( P = 0.0001 female, and P = 0.0011 male) in H over both L and M ( P < 0.0001 for each). In males, GEO of the Poly I:C group was higher than that of the ITP group ( P = 0.0374), with no difference observed among females over all days. By using a FCM algorithm for population comparisons of fluorescence to investigate binding within H, the T(x) scores indicated higher fluorescence in control and Poly I:C groups over ITP ( P = 0.0001). Unlike chickens, Gallus gallus , which express CD4 primarily on L, kestrels bound the commercial antibody primarily within the gated H subpopulation, suggesting an immunophenotypic difference between taxa, despite a ~60% identity of Falco CD4 amino acid sequences with chicken CD4. The emergent cell subset within the gated H presented dendritic-like cell (DLC) morphological and functional properties, apparently serving as an effector cell. This study adds interpretive context to ecological investigations of infection and of potential immunomodulation by emerging compounds, whereby the early innate responses are mediated by the various cell subsets serving as useful quantitative markers of immunological condition. Data showed that dietary exposure to ITP was immunosuppressive for male and female kestrels over the course of the experiment, reducing DLC frequency compared to the Poly I:C controls. Heterophils and DLC were important in facilitating innate immunological responses.

Environment International↗

Parallel Fortran-MPI software for numerical inversion of the Laplace transform and its application to oscillatory water levels in groundwater environments

A parallel Fortran-MPI (Message Passing Interface) software for numerical inversion of the Laplace transform based on a Fourier series method is developed to meet the need of solving intensive computational problems involving oscillatory water level's response to hydraulic tests in a groundwater environment. The software is a parallel version of ACM (The Association for Computing Machinery) Transactions on Mathematical Software (TOMS) Algorithm 796. Running 38 test examples indicated that implementation of MPI techniques with distributed memory architecture speedups the processing and improves the efficiency. Applications to oscillatory water levels in a well during aquifer tests are presented to illustrate how this package can be applied to solve complicated environmental problems involved in differential and integral equations. The package is free and is easy to use for people with little or no previous experience in using MPI but who wish to get off to a quick start in parallel computing. ?? 2004 Elsevier Ltd. All rights reserved.

Environmental Modelling and Software↗

A computer program for uncertainty analysis integrating regression and Bayesian methods

This work develops a new functionality in UCODE_2014 to evaluate Bayesian credible intervals using the Markov Chain Monte Carlo (MCMC) method. The MCMC capability in UCODE_2014 is based on the FORTRAN version of the differential evolution adaptive Metropolis (DREAM) algorithm of Vrugt et al. (2009), which estimates the posterior probability density function of model parameters in high-dimensional and multimodal sampling problems. The UCODE MCMC capability provides eleven prior probability distributions and three ways to initialize the sampling process. It evaluates parametric and predictive uncertainties and it has parallel computing capability based on multiple chains to accelerate the sampling process. This paper tests and demonstrates the MCMC capability using a 10-dimensional multimodal mathematical function, a 100-dimensional Gaussian function, and a groundwater reactive transport model. The use of the MCMC capability is made straightforward and flexible by adopting the JUPITER API protocol. With the new MCMC capability, UCODE_2014 can be used to calculate three types of uncertainty intervals, which all can account for prior information: (1) linear confidence intervals which require linearity and Gaussian error assumptions and typically 10s–100s of highly parallelizable model runs after optimization, (2) nonlinear confidence intervals which require a smooth objective function surface and Gaussian observation error assumptions and typically 100s–1,000s of partially parallelizable model runs after optimization, and (3) MCMC Bayesian credible intervals which require few assumptions and commonly 10,000s–100,000s or more partially parallelizable model runs. Ready access allows users to select methods best suited to their work, and to compare methods in many circumstances.

Environmental Modelling and Software↗

The U. S. Geological Survey National Hydrologic Model infrastructure: Rationale, description, and application of a watershed-scale model for the conterminous United States

The ability to effectively manage water resources to meet present and future human and environmental needs is essential. Such an ability necessitates a comprehensive understanding of hydrologic processes that affect streamflow at a watershed scale. In the United States, water-resources management at scales ranging from local to national can benefit from a nationally consistent, process-based watershed modeling capability to provide the requisite understanding. The National Hydrologic Model (NHM) infrastructure, which was developed by the U.S. Geological Survey to support coordinated, comprehensive, and consistent hydrologic modeling at multiple scales for the conterminous United States, provides this essential capability. NHM-based applications provide information to enable more effective water-resources planning and management, fill knowledge gaps in ungaged areas, and support basic scientific inquiry. In the future, as process algorithms and data sets improve, the NHM infrastructure will continue to evolve to better support the nation's water-resources research and management needs.

Conterminous United States↗

A model-independent tool for evolutionary constrained multi-objective optimization under uncertainty

An open-source tool has been developed to facilitate constrained single- and multi-objective optimization under uncertainty (CMOU) analyses. The tool uses the well-known PEST interface protocols to communicate with the underlying forward simulation, making it non-intrusive. The tool contains a built-in parallel run manager to make use of heterogeneous and distributed computing resources. Several popular and well-known evolutionary algorithms are implemented and can be combined with a range of approaches to represent uncertainty in model-derived constraint/objective values. These attributes serve to address the current barrier to adopt advanced CMOU analyses for a wide range of decision-support problems across the environmental modeling spectrum. We demonstrate the capabilities of the CMOU tool on a well-known analytical benchmark problem that we augmented to include uncertainty, as well as on a synthetic density-dependent coastal groundwater management benchmark problem. Both demonstrations highlight the importance of explicitly accounting for uncertainty to convey risk and reliability in pareto-optimal design.

Environmental Modelling & Software↗

A machine learning approach to predicting equilibrium ripple wavelength

Sand ripples are geomorphic features on the seafloor that affect bottom boundary layer dynamics including wave attenuation and sediment transport. We present a new equilibrium ripple predictor using a machine learning approach that outputs a probability distribution of wave-generated equilibrium wavelengths and statistics including an estimate of ripple height, the most probable ripple wavelength, and sediment and flow parameterizations. The Bayesian Optimal Model System (BOMS) is an ensemble machine learning system that combines two machine learning algorithms and two deterministic empirical ripple predictors with a Bayesian meta-learner to produce probabilistic wave-generated equilibrium ripple wavelength estimates in sandy locations. A ten-fold cross validation of BOMS resulted in an adjusted R-squared value of 0.93 and an average root mean square error (RMSE) of 8.0 cm. During both cross validation and testing on three unique field datasets, BOMS provided more accurate wavelength predictions than each individual base model and other common ripple predictors.

Environmental Modeling and Software↗

Climate matching with the climatchR R package

Climate matching allows comparisons of climatic conditions between different locations to understand location and species range climatic suitability. The approach may be used as part of horizon scanning exercises such as those conducted for invasive species. We implemented the CLIMATCH algorithm into an R package, climatchR . The package allows automated and scripted climate matching exercises across all steps from downloading data to summarizing species climate matches. We also show how climatchR may be used with high-throughput computing to process many species. For example, we were able to calculate climate scores for over 8,000 species in less than 3 days using this package. This automation allows high-throughput processing of species data, a new development for improving the efficiency and speed of climate matching and horizon scanning.

Environmental Software & Modeling↗

Investigating hydrologic alteration in the Pearl and Pascagoula River basins using rule-based model trees

Anthropogenic hydrologic alteration threatens the health of riverine ecosystems. Machine learning algorithms that employ the use of model trees to predict hydrologic alteration are underrepresented in related literature. This study assesses hydrologic alteration in the Pearl and Pascagoula River basins using modeled daily streamflow . Hydrologic alteration was determined by hypothesis testing and the computation of the net change across 60 years. Cubist models were developed for both basins to predict hydrologic alteration and to identify important basin characteristics. Results from net change and the hypothesis test indicated the basins were essentially identical with respect to the amount of hydrologic alteration. Cubist models for the basins successfully made accurate predictions of hydrologic alteration and demonstrated that the importance of basin geomorphology and land cover on alteration differed in both basins. The results of the study demonstrate the feasibility of model trees in assessing hydrologic alteration.

Mississippi↗

Bayesian estimation of magma supply, storage, and eruption rates using a multiphysical volcano model: Kīlauea Volcano, 2000–2012

Estimating rates of magma supply to the world's volcanoes remains one of the most fundamental aims of volcanology. Yet, supply rates can be difficult to estimate even at well-monitored volcanoes, in part because observations are noisy and are usually considered independently rather than as part of a holistic system. In this work we demonstrate a technique for probabilistically estimating time-variable rates of magma supply to a volcano through probabilistic constraint on storage and eruption rates. This approach utilizes Bayesian joint inversion of diverse datasets using predictions from a multiphysical volcano model, and independent prior information derived from previous geophysical, geochemical, and geological studies. The solution to the inverse problem takes the form of a probability density function which takes into account uncertainties in observations and prior information, and which we sample using a Markov chain Monte Carlo algorithm. Applying the technique to Kīlauea Volcano, we develop a model which relates magma flow rates with deformation of the volcano's surface, sulfur dioxide emission rates, lava flow field volumes, and composition of the volcano's basaltic magma. This model accounts for effects and processes mostly neglected in previous supply rate estimates at Kīlauea, including magma compressibility, loss of sulfur to the hydrothermal system, and potential magma storage in the volcano's deep rift zones. We jointly invert data and prior information to estimate rates of supply, storage, and eruption during three recent quasi-steady-state periods at the volcano. Results shed new light on the time-variability of magma supply to Kīlauea, which we find to have increased by 35&ndash;100% between 2001 and 2006 (from 0.11&ndash;0.17 to 0.18&ndash;0.28 km 3 /yr), before subsequently decreasing to 0.08&ndash;0.12 km 3 /yr by 2012. Changes in supply rate directly impact hazard at the volcano, and were largely responsible for an increase in eruption rate of 60&ndash;150% between 2001 and 2006, and subsequent decline by as much as 60% by 2012. We also demonstrate the occurrence of temporal changes in the proportion of Kīlauea's magma supply that is stored versus erupted, with the supply &ldquo;surge&rdquo; in 2006 associated with increased accumulation of magma at the summit. Finally, we are able to place some constraints on sulfur concentrations in Kīlauea magma and the scrubbing of sulfur by the volcano's hydrothermal system. Multiphysical, Bayesian constraint on magma flow rates may be used to monitor evolving volcanic hazard not just at Kīlauea but at other volcanoes around the world.

Hawai'i↗

Using active source seismology to image the Palos Verdes Fault damage zone as a function of distance, depth, and geology

Fault damage zones provide a window into the non-elastic processes of an earthquake. Geological and seismic tomography methods have been unable to measure damage zones at depth with sufficient spatial sampling to evaluate the relative influence of depth, distance, and lithological variations. Here, we identify and analyze the damage zone of the Palos Verdes Fault offshore southern California using two 3D seismic reflection datasets. We apply a novel algorithm to identify discontinuities attributed to faults and fractures in large seismic volumes and examine the spatial distribution of fault damage in sedimentary rock surrounding the Palos Verdes Fault. Our results show that damage through fracturing is most concentrated around mapped faults and decays exponentially to a distance of ∼2 km, where fracturing reaches a clearly defined and relatively undamaged background for all examined depths and lithologies (450 m to 2.2 km). This decrease in fracturing with distance from the central fault strand exhibits similar functional form to outcrop studies. However, here we extend analysis to distances seldom accessible (∼10 km lateral distance). Separating the data by geologic units we find that the damage decay and background level differs for each unit, with the older and deeper units having higher levels of background fracturing and shallower exponential decays of fracturing with distance from the fault. Surprisingly, these differences in damage decay and background level trade-off result in a consistent damage zone width regardless of lithology or depth. We find that the damage zone has similar decay trends on both sides of the fault. When examining the damage zone at shorter (4 km vs 17 km) along strike distances, the damage zone has a more complex decay trend and at least two strands are resolvable.

California↗

Snag dynamics and surface fuel loads in the Sierra Nevada: Predicting the impact of the 2012–2016 drought

Forest die-backs linked to extreme droughts are expected to increase as the climate dries and warms. An example is the 2012-2016 hotter drought in California that induced widespread tree mortality in the Sierra Nevada, California. The sudden increase in snags (i.e., standing dead trees) raised immediate concerns about their impact on wildfire hazard and longer-term questions about their impact on ecosystem structure and function. We quantified the likely progression of snag fall and fuel succession following the recent extensive mortality event in the southern Sierra Nevada mixed conifer forest. Our results used data from a long-term demography study to project trends in surface fuel loads at three study sites in Yosemite and Sequoia Kings Canyon National Parks. In the short term (2017-2021), fine woody debris and litter + duff significantly increased across all three sites (>145% and >55%, respectively); coarse woody debris increased significantly at one site (48.6%); and total fuel loads increased significantly at two of the three sites (38% and 69%). Snag longevity increased with size, with the relationship varying by species. Yellow pine was a notable outlier: size played a small role in influencing its fall rates. Overall, species-specific snag fall rates in the southern Sierra Nevada were 20% to 40% slower than previously reported. By 2040, projected median cumulative inputs of biomass from future snag fall range from 49.4 Mg ha-1 to 136.1 Mg ha-1across our three sites, which exceeds the amounts currently present (47.17-89.97 Mg ha-1) and is well above estimates of historical coarse woody debris amounts in the Sierra Nevada (17.7 Mg ha -1). These results provide a robust empirical basis to refine the snag fall algorithm in vegetation simulation models. Options to manage the impact of extreme number of snags and their large surface combustible biomass include salvage operations and prescribed burning, with both methods having operational, financial, and legal limitations that need to be considered.

Forest Ecology and Management↗

Machine learning for predicting soil classes in three semi-arid landscapes

Mapping the spatial distribution of soil taxonomic classes is important for informing soil use and management decisions. Digital soil mapping (DSM) can quantitatively predict the spatial distribution of soil taxonomic classes. Key components of DSM are the method and the set of environmental covariates used to predict soil classes. Machine learning is a general term for a broad set of statistical modeling techniques. Many different machine learning models have been applied in the literature and there are different approaches for selecting covariates for DSM. However, there is little guidance as to which, if any, machine learning model and covariate set might be optimal for predicting soil classes across different landscapes. Our objective was to compare multiple machine learning models and covariate sets for predicting soil taxonomic classes at three geographically distinct areas in the semi-arid western United States of America (southern New Mexico, southwestern Utah, and northeastern Wyoming). All three areas were the focus of digital soil mapping studies. Sampling sites at each study area were selected using conditioned Latin hypercube sampling (cLHS). We compared models that had been used in other DSM studies, including clustering algorithms, discriminant analysis, multinomial logistic regression, neural networks, tree based methods, and support vector machine classifiers. Tested machine learning models were divided into three groups based on model complexity: simple, moderate, and complex. We also compared environmental covariates derived from digital elevation models and Landsat imagery that were divided into three different sets: 1) covariates selected a priori by soil scientists familiar with each area and used as input into cLHS, 2) the covariates in set 1 plus 113 additional covariates, and 3) covariates selected using recursive feature elimination. Overall, complex models were consistently more accurate than simple or moderately complex models. Random forests (RF) using covariates selected via recursive feature elimination was consistently the most accurate, or was among the most accurate, classifiers between study areas and between covariate sets within each study area. We recommend that for soil taxonomic class prediction, complex models and covariates selected by recursive feature elimination be used. Overall classification accuracy in each study area was largely dependent upon the number of soil taxonomic classes and the frequency distribution of pedon observations between taxonomic classes. Individual subgroup class accuracy was generally dependent upon the number of soil pedon observations in each taxonomic class. The number of soil classes is related to the inherent variability of a given area. The imbalance of soil pedon observations between classes is likely related to cLHS. Imbalanced frequency distributions of soil pedon observations between classes must be addressed to improve model accuracy. Solutions include increasing the number of soil pedon observations in classes with few observations or decreasing the number of classes. Spatial predictions using the most accurate models generally agree with expected soil–landscape relationships. Spatial prediction uncertainty was lowest in areas of relatively low relief for each study area.

New Mexico, Utah, Wyoming↗

POLARIS: A 30-meter probabilistic soil series map of the contiguous United States

A new complete map of soil series probabilities has been produced for the contiguous United States at a 30 m spatial resolution. This innovative database, named POLARIS, is constructed using available high-resolution geospatial environmental data and a state-of-the-art machine learning algorithm (DSMART-HPC) to remap the Soil Survey Geographic (SSURGO) database. This 9 billion grid cell database is possible using available high performance computing resources. POLARIS provides a spatially continuous, internally consistent, quantitative prediction of soil series. It offers potential solutions to the primary weaknesses in SSURGO: 1) unmapped areas are gap-filled using survey data from the surrounding regions, 2) the artificial discontinuities at political boundaries are removed, and 3) the use of high resolution environmental covariate data leads to a spatial disaggregation of the coarse polygons. The geospatial environmental covariates that have the largest role in assembling POLARIS over the contiguous United States (CONUS) are fine-scale (30 m) elevation data and coarse-scale (~ 2 km) estimates of the geographic distribution of uranium, thorium, and potassium. A preliminary validation of POLARIS using the NRCS National Soil Information System (NASIS) database shows variable performance over CONUS. In general, the best performance is obtained at grid cells where DSMART-HPC is most able to reduce the chance of misclassification. The important role of environmental covariates in limiting prediction uncertainty suggests including additional covariates is pivotal to improving POLARIS' accuracy. This database has the potential to improve the modeling of biogeochemical, water, and energy cycles in environmental models; enhance availability of data for precision agriculture; and assist hydrologic monitoring and forecasting to ensure food and water security.

Geoderma↗