Geology ReportsSearch

SEARCH · Geology Reports

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Application of AFINCH as a tool for evaluating the effects of streamflow-gaging-network size and composition on the accuracy and precision of streamflow estimates at ungaged locations in the southeast Lake Michigan hydrologic subregion

Bootstrapping techniques employing random subsampling were used with the AFINCH (Analysis of Flows In Networks of CHannels) model to gain insights into the effects of variation in streamflow-gaging-network size and composition on the accuracy and precision of streamflow estimates at ungaged locations in the 0405 (Southeast Lake Michigan) hydrologic subregion. AFINCH uses stepwise-regression techniques to estimate monthly water yields from catchments based on geospatial-climate and land-cover data in combination with available streamflow and water-use data. Calculations are performed on a hydrologic-subregion scale for each catchment and stream reach contained in a National Hydrography Dataset Plus (NHDPlus) subregion. Water yields from contributing catchments are multiplied by catchment areas and resulting flow values are accumulated to compute streamflows in stream reaches which are referred to as flow lines. AFINCH imposes constraints on water yields to ensure that observed streamflows are conserved at gaged locations. Data from the 0405 hydrologic subregion (referred to as Southeast Lake Michigan) were used for the analyses. Daily streamflow data were measured in the subregion for 1 or more years at a total of 75 streamflow-gaging stations during the analysis period which spanned water years 1971–2003. The number of streamflow gages in operation each year during the analysis period ranged from 42 to 56 and averaged 47. Six sets (one set for each censoring level), each composed of 30 random subsets of the 75 streamflow gages, were created by censoring (removing) approximately 10, 20, 30, 40, 50, and 75 percent of the streamflow gages (the actual percentage of operating streamflow gages censored for each set varied from year to year, and within the year from subset to subset, but averaged approximately the indicated percentages). Streamflow estimates for six flow lines each were aggregated by censoring level, and results were analyzed to assess (a) how the size and composition of the streamflow-gaging network affected the average apparent errors and variability of the estimated flows and (b) whether results for certain months were more variable than for others. The six flow lines were categorized into one of three types depending upon their network topology and position relative to operating streamflow-gaging stations. Statistical analysis of the model results indicates that (1) less precise (that is, more variable) estimates resulted from smaller streamflow-gaging networks as compared to larger streamflow-gaging networks, (2) precision of AFINCH flow estimates at an ungaged flow line is improved by operation of one or more streamflow gages upstream and (or) downstream in the enclosing basin, (3) no consistent seasonal trend in estimate variability was evident, and (4) flow lines from ungaged basins appeared to exhibit the smallest absolute apparent percent errors (APEs) and smallest changes in average APE as a function of increasing censoring level. The counterintuitive results described in item (4) above likely reflect both the nature of the base-streamflow estimate from which the errors were computed and insensitivity in the average model-derived estimates to changes in the streamflow-gaging-network size and composition. Another analysis demonstrated that errors for flow lines in ungaged basins have the potential to be much larger than indicated by their APEs if measured relative to their true (but unknown) flows. “Missing gage” analyses, based on examination of censoring subset results where the streamflow gage of interest was omitted from the calibration data set, were done to better understand the true error characteristics for ungaged flow lines as a function of network size. Results examined for 2 water years indicated that the probability of computing a monthly streamflow estimate within 10 percent of the true value with AFINCH decreased from greater than 0.9 at about a 10-percent network-censoring level to less than 0.6 as the censoring level approached 75 percent. In addition, estimates for typically dry months tended to be characterized by larger percent errors than typically wetter months.

Indiana, Michigan

Examination of flood characteristics at selected streamgages in the Meramec River Basin, eastern Missouri, December 2015–January 2016

Overview Heavy rainfall resulted in major flooding in the Meramec River Basin in eastern Missouri during late December 2015 through early January 2016. Cumulative rainfall from December 14 to 29, 2015, ranged from 7.6 to 12.3 inches at selected precipitation stations in the basin with flooding driven by the heaviest precipitation (3.9–9.7 inches) between December 27 and 29, 2015. Financial losses from flooding included damage to homes and other structures, damage to roads, and debris removal. Eight of 11 counties in the basin were declared a Federal Disaster Area. The U.S. Geological Survey (USGS), in cooperation with the U.S. Army Corps of Engineers and St. Louis Metropolitan Sewer District, operates multiple streamgages along the Meramec River and its primary tributaries including the Bourbeuse River and Big River. The period of record for streamflow at streamgages in the basin included in this report ranges from 24 to 102 years. Instrumentation in a streamgage shelter automatically makes observations of stage using a variety of methods (submersible pressure transducer, non-submersible pressure transducer, or non-contact radar). These observations are recorded autonomously at a predetermined programmed frequency (typically either 15 or 30 minutes) dependent on drainage-area size and concomitant flashiness of the stream. Although stage data are important, streamflow data are equally or more important for streamflow forecasting, water-quality constituent loads computation, flood-frequency analysis, and flood mitigation planning. Streamflows are computed from recorded stage data using an empirically determined relation between stage and streamflow termed a “rating.” Development and verification of the rating requires periodic onsite discrete measurements of streamflow throughout time and over the range of stages to define local hydraulic conditions. The purpose of this report is to examine characteristics of flooding that occurred in the Meramec River Basin in December 2015–January 2016 including peak stages, peak streamflows, and the flood-frequency statistics associated with the peak flows. A comparison between the December 2015–January 2016 flood and a similar flood in December 1982 in the Meramec River Basin also is included.

Missouri

Trend analysis and selected summary statistics of annual mean streamflow for 38 selected long-term U.S. Geological Survey streamgages in Texas, water years 1916-2012

In 2013, the U.S. Geological Survey (USGS) operated more than 500 continuous streamgages (streamflow-gaging stations) in Texas. In cooperation with the Texas Water Development Board, the USGS evaluated mean annual streamflow data for 38 selected streamgages that were active as of water year 2012. The 38 streamgages have annual mean streamflow data considered natural and unregulated. Collected annual mean streamflow data for a single streamgage ranged from 49 to 97 cumulative years. The nonparametric Kendall’s tau statistical test was used to detect monotonic trends in annual mean streamflow over time. The monotonic trend analysis detected 2 statistically significant upward trends (0.01 one-tail significance), 1 statistically significant downward trend (0.01 one-tail significance level), and 35 instances of no statistically significant trend (0.02 two-tailed significance level). The Theil slope estimate of a regression slope of annual mean streamflow with time was computed for the three stations where trends in streamflow were detected: 2 increasing Theil slopes were measured (+0.40 and +2.72 cubic feet per second per year, respectively), and 1 decreasing Theil slope (–0.24 cubic feet per second per year) was measured. Selected summary statistics (L-moments) and estimates of respective sampling variances were computed for the 35 streamgages lacking statistically significant trends. From the L-moments and estimated sampling variances, weighted means or regional values were computed for each L-moment. An example application is included demonstrating how the L-moments could be used to evaluate the magnitude and frequency of annual mean streamflow.

Texas

Palynological applications of principal component and cluster analyses

Two multivariate statistical methods are suggested to help describe patterns in pollen data that result from changes in the relative frequencies of pollen types produced by past climatic and environmental variations. These methods, based on a geometric model, compare samples by use of the product-moment correlation coefficient computed from data subjected to a centering transformation. If there are m samples and n pollen types, then the data can be regarded as a set of m points in an n -dimensional space. The first method, cluster analysis, produces a dendrograph or clustering tree in which samples are grouped with other samples on the basis of their similarity to each other. The second method, principal component analysis, produces a set of variates that arc linear combinations of the pollen samples, are uncorrelated with each other, and best describe the data using a minimum number of dimensions. This method is useful in reducing the dimensionality of data sets. A further transformation known as varimax rotation acts on a subset of the principal components to make them easier to interpret. Both methods offer the advantages of reproducibility of results and speed in pattern description. Once the patterns in the data have been described, however, they must be interpreted by the palynologist. An application of the methods in palynology is shown by using data from Osgood Swamp, Calif.

Journal of Research of the U.S. Geological Survey

A regional classification of the effectiveness of depressional wetlands at mitigating nitrogen transport to surface waters in the Northern Atlantic Coastal Plain

Nitrogen from nonpoint sources contributes to eutrophication, hypoxia, and related ecological degradation in Atlantic Coastal Plain streams and adjacent coastal estuaries such as Chesapeake Bay and Pamlico Sound. Although denitrification in depressional (non-riparian) wetlands common to the Coastal Plain can be a significant landscape sink for nitrogen, the effectiveness of individual wetlands at removing nitrogen varies substantially due to varying hydrogeologic, geochemical, and other landscape conditions, which are often poorly or inconsistently mapped over large areas. A geographic model describing the spatial variability in the likely effectiveness of depressional wetlands in watershed uplands at mitigating nitrogen transport from nonpoint sources to surface waters was constructed for the Northern Atlantic Coastal Plain (NACP), from North Carolina through New Jersey. Geographic and statistical techniques were used to develop the model. Available medium-resolution (1:100,000-scale) stream hydrography was used to define 33,799 individual watershed catchments in the study area. Sixteen landscape metrics relevant to the occurrence of depressional wetlands and their effectiveness as nitrogen sinks were defined for each catchment, based primarily on available topographic and soils data. Cluster analysis was used to aggregate the 33,799 catchments into eight wetland landscape regions (WLRs) based on the value of three principal components computed for the 16 original landscape metrics. Significant differences in topography, soil, and land cover among the eight WLRs demonstrate the effectiveness of the clustering technique. Results were used to interpret the relative likelihood of depressional wetlands in each WLR and their likely effectiveness at mitigating nitrogen transport from upland source areas to surface waters. The potential effectiveness of depressional wetlands at mitigating nitrogen transport varies substantially over different parts of the NACP. Depressional wetlands are common in three WLRs covering 32 percent of the area, and have a relatively high potential to mitigate nitrogen transport from nonpoint sources. Conversely, 37 percent of the study area includes rolling hills with relatively high slope and relief, and little likelihood of depressional wetlands. The remainder of the Coastal Plain includes relatively flat watersheds with moderate to low relative likelihood of nitrogen mitigation. The delineation of WLRs in this model should be useful for targeting wetland conservation or restoration efforts, and for estimating the effects of depressional wetlands on the regional nitrogen budget, but should be considered in light of limitations and assumptions inherent in the model.

Scientific Investigations Report

Variability of differences between two approaches for determining ground-water discharge and pumpage, including effects of time trends, Lower Arkansas River Basin, southeastern Colorado, 1998-2002

In the mid-1990s, the Colorado Division of Water Resources (CDWR) adopted rules governing measurement of tributary ground-water pumpage for the Arkansas River Basin. The rules allowed ground-water pumpage to be determined using one of two approaches?power conversion coefficient (PCC) or totalizing flowmeters (TFM). In addition, the rules allowed a PCC to be applied to the electrical power usage up to 4 years in the future to estimate ground-water pumpage. As a result of concerns about potential errors in applying the PCC approach forward in time, a study was done by the U.S. Geological Survey, in cooperation with CDWR and Colorado Water Conservation Board, to evaluate the variability in differences in pumpage between the two approaches, including the effects of time trends. This report compared measured ground-water pumpage using TFMs to computed ground-water pumpage using PCCs by developing statistical models of relations between explanatory variables, such as site, time, and pumping water level, and dependent variables, which are based on discharge, PCC, and pumpage. When differences in pumpage (diffP) were computed using PCC measurements and power consumption for the same year (1998-2002), the median diffP, depending on the year, ranged from +0.1 to -2.9 percent; the median diffP for the entire period was -1.5 percent. However, when diffP was computed using PCC measurements applied to the next year's power consumption, the median diffP was -0.3 percent; and when PCC measurements were applied 2, 3, or 4 years into the future, median diffPs were +1.8 percent for a 2-year forward lag and +5.3 percent for a 4-year forward lag, indicating that pumpage computed with the PCC approach, as generally applied under the ground-water pumpage measurement rules by CDWR, tended to overestimate pumpage as compared to pumpage using TFMs when PCC measurement was applied to future years of measured power consumption. Analyses were done to better understand the causes of the time trend; an estimate of the overall trend with time (uncorrected for pumping water-level changes) yielded a trend of about 2.2 percent per lag year for diffP. A separate analysis that incorporated a surface-water diversion term in the statistical model rendered the time-trend term insignificant, indicating that the time trend in the models served as a surrogate for other variables, some of which reflect underlying hydrologic conditions. A more precise explanation of the potential causes of the time trend was not obtained with the available data. However, the model results with the surface-water diversion term indicate that much of the trend of 2.2 percent per lag year in diffP resulted from applying a PCC to estimate pumpage under hydrologic conditions different from those under which the PCC was measured. Although there is no evidence to conclude that the upward time trend determined in the data for this 5-year period would hold in the future, historical static ground-water levels in the study area generally have exhibited small variations over multidecadal time scales. Therefore, the approximately 2 percent per lag year trend determined in these data is expected to be a reasonable guideline for estimating potential errors in the PCC approach resulting from temporally varying hydrologic conditions between time of PCC measurement and pumpage estimation. Comparisons also were made between total, or aggregated, pumpage for a network of wells as computed by the PCC approach and the TFM approach. For 100 wells and a lag of 4 years between PCC measurement and pumpage estimation, there was a 95-percent probability that the difference between total network pumpage measured by the PCC approach and that measured using a TFM would be between 5.2 and 14.4 percent. These estimates were based on a bias of 2.2 percent per lag year estimated for the period 1998-2002 during which hydrologic conditions were known to have changed. Using the same assumptions, the estimated d

Scientific Investigations Report

Global cropland-extent product at 30-m resolution (GCEP30) derived from Landsat satellite time-series data for the year 2015 using multiple machine-learning algorithms on Google Earth Engine cloud

Executive Summary Global food and water security analysis and management require precise and accurate global cropland-extent maps. Existing maps have limitations, in that they are (1) mapped using coarse-resolution remote-sensing data, resulting in the lack of precise mapping location of croplands and their accuracies; (2) derived by collecting and collating national statistical data that are often subjective, leading to substantial uncertainties in cropland-area estimates, as well as their locations; and (3) extracted from one or more classes of a land use–land cover product in which cropland classes are not the focus of mapping, leading to their mixing with other classes and creating significant errors of omission and commission. These limitations can be overcome by producing high-resolution cropland-extent maps using satellite-sensor data, such as Landsat 30-m resolution or higher. The most fundamental cropland product is the high-resolution cropland-extent map because all higher level cropland products, such as crop-watering method (that is, whether crops are irrigated or rainfed), crop types, cropping intensities, cropland fallows, crop productivity, and crop-water productivity, are dependent on a precise and accurate cropland-extent product. Given these realities, the overarching goal of this study was to produce a Landsat satellite-derived global cropland-extent product at 30-m resolution. The work, which involved a paradigm shift in how global cropland-extent maps are produced, involved the following five key steps: (1) petabyte-scale computing that involved multiyear, 8- to 16-day, time-series Landsat 30-m resolution data for the global land surface; (2) composition of analysis-ready data (ARD) cubes; (3) creation of a large global-reference data hub for machine learning; (4) use of multiple machine-learning algorithms (MLAs) by writing software and computing in the cloud; and (5) Google Earth Engine (GEE) cloud computing. The five key steps involved nine distinct phases. First, the world was segmented into 74 agroecological zones (AEZs). Second, Landsat 8- to 16-day data were used to time-composite 10-band (blue, green, red, near-infrared, short-wave infrared band 1, short-wave infrared band 2, thermal infrared, enhanced vegetation index, normalized difference water index, and normalized difference vegetation index) Landsat 30-m resolution data cubes for every 2- to 4-month time period during 3- to 4-year periods (stated as nominal-year 2015 or, simply, 2015), along with two additional 30-m resolution bands (Shuttle Radar Topography Mission elevation, and slope) in each of the 74 AEZs. Third, more than 100,000 reference-training data samples were collected using ground data (some of which were collected using a mobile application), as well as submeter- to 5-m-resolution, very high-resolution imagery sourced from other reliable sources. Fourth, reference-training data were used to create a knowledge base for separating cropland from noncropland. Fifth, MLAs such as the pixel-based supervised random forest and support-vector machines were written on the GEE using Python and JavaScript. Sixth, object-based recursive hierarchical segmentation algorithm was used, in addition to MLAs, to overcome uncertainties. Seventh, MLAs used the knowledge base to classify and separate cropland from noncropland. Eighth, accuracy assessment was conducted by generating error matrices for each of the 74 AEZs using 19,171 independent validation-data samples. Ninth, cropland areas were computed for all countries of the world and compared with United Nation’s (UN’s) Food and Agricultural Organization (FAO) and other national statistics. The outcome was a Landsat-derived global cropland-extent product at 30-m resolution (GCEP30), which has an overall accuracy of 91.7 percent. For the cropland class, producer’s accuracy was 83.4 percent, and user’s accuracy was 78.3 percent. GCEP30 calculated (using direct pixel count) the global net-cropland area (GNCA) for the year 2015 as 1.873 billion hectares (~12.6 percent of the Earth’s terrestrial area). The continental cropland distribution as a percentage of GNCA was Asia, 33 percent; Europe, 25.5 percent; Africa, 16.7 percent; North America, 14.4 percent; South America, 8.1 percent; and Australia and Oceania, 2.4 percent. The worldwide cropland areas in GCEP30 for 2015 were higher by 236 to 299 million hectares (Mha) compared to national statistics reported elsewhere for the same year (for example, in Food and Agriculture Organization’s corporate statistical database [FAOSTAT] and in the monthly irrigated and rainfed crop areas [MIRCA] database). The global cropland area reported for 2015 increased by 344 Mha (22.5 percent), compared to the year 2000. During the same period (2000–2015), the world’s population increased by 20 percent. Whereas some of these areal increases are real increases in cropland areas, others are due to the types of data, methods, and approaches used. Using the highest known resolution (compared to previous coarse-resolution global products) enabled this study to capture fragmented croplands. Coarse-resolution data compute areas on the basis of subpixels, which, for a large proportion of certain land use–land cover classes, will show only a certain percentage of the total pixel area as actual area. Subpixel areas can lead to substantial uncertainties in area computation, as determining the exact fraction of cropland areas within a coarse-resolution pixel is resource intensive and subject to errors. Other innovations in GCEP30 include reference-data hubs, machine learning, and cloud computing. Cropland areas in 214 countries, territories, departments, and regions were calculated for the year 2015 using GCEP30, on the basis of UN’s global administrative unit layers (GAUL) boundaries. The 10 leading countries in terms of cropland area (as a percentage of the GNCA) were India (9.6 percent), United States (8.95 percent), China (8.82 percent), Russia (8.32 percent), Brazil (3.42 percent), Ukraine (2.32 percent), Canada (2.29 percent), Argentina (2.05 percent), Indonesia (2 percent), and Nigeria (1.91 percent). Together, these 10 countries occupy 50 percent of the global cropland, and they have 52 percent of the global population. Their combined cropland area increased by 2 percent between 2000 and 2015, compared to the substantial increase in population of 517 million (15.5 percent). Together, India, United States, China, and Russia encompass 36 percent of the total area. In the United States and Canada, from 2000 to 2015, cropland decreased by about 2 percent, whereas their populations increased by 14 and 13 percent, respectively. The additional food requirements in these 10 countries, which are caused by increased populations, as well as increasing nutritional demands, are met by production increases in existing cropland or through virtual food trade, or both. More than 18 countries, territories, departments, or regions had 60 percent or more of their geographic area as cropland: Republic of Moldova, San Marino, and Hungary had more than 80 percent of the country’s area as cropland; Denmark, Ukraine, Ireland, and Bangladesh, 70 to 80 percent; and Uruguay, Netherlands, United Kingdom, Spain, Lithuania, Poland, Gaza Strip, Czechia, Italy, India, and Azerbaijan, 60 to 70 percent. Europe and South Asia can be considered agricultural capitals of the world, on the basis of their percentages of geographic area as cropland. United States, China, and Russia, which all have high cropland areas, are ranked second, third, and fourth in the world; India is ranked first. However, the amount of cropland as a percentage of the country’s geographic area is relatively very low for United States (18.3 percent), China (17.7 percent), and Russia (9.5 percent), whereas it is 60.5 percent for India. Most African and South American countries, territories, departments, or regions have less than 15 percent of their geographic area as cropland. China and India together house 36 percent of the world’s population; however, between 2000 and 2015, the amount of China’s cropland area fell by 18.9 percent, owing to urban expansion and the abandonment of farmlands caused by demographic changes (that is, the movement of population from villages to cities). In contrast, China’s population grew by 10 percent. The amount of India’s cropland increased by 8.5 percent, whereas its population grew by 20 percent. This study showed that, out of the 10 leading cropland countries, Ukraine, Nigeria, Russia, and Indonesia showed an 18 to 31 percent increase in cropland areas, on the basis of GCEP30 by the year 2015, compared to 2000. Nigeria’s cropland area increased by 25 percent, and its population increased by 31 percent in the same period. In these countries, food security is maintained by cropland expansion, productivity increases, and virtual food trade. Nevertheless, this trend of increasing net-cropland area and productivity will likely become difficult to maintain, owing to diminishing arable lands and plateauing of 50 years of continual yield increases, requiring policymakers to explore novel and data-supported approaches to solving future food security issues. The GCEP30 product, which can be browsed at full resolution at www.croplands.org , has been released for public download and use through U.S. Geological Survey (USGS)–National Aeronautics and Space Administration (NASA) Land Processes Distributed Active Archive Center (see https://lpdaac.usgs.gov/news/release-of-gfsad-30-meter-cropland-extent-products/ ).

Professional Paper

Application of a process-based shallow landslide hazard model over a broad area in Central Italy

Process-based models are widely used for rainfall-induced shallow landslide forecasting. Previous studies have successfully applied the U.S. Geological Survey’s Transient Rainfall Infiltration and Grid-Based Regional Slope-Stability (TRIGRS) model (Baum et al. 2002 ) to compute infiltration-driven changes in the hillslopes’ factor of safety on small scales (i.e., tens of square kilometers). Soil data input for such models are difficult to obtain across larger regions. This work describes a novel methodology for the application of TRIGRS over broad areas with relatively uniform hydrogeological properties. The study area is a 550-km 2 region in Central Italy covered by post-orogenic Quaternary sediments. Due to the lack of field data, we assigned mechanical and hydrological property values through a statistical analysis based on literature review of soils matching the local lithologies. We calibrated the model using rainfall data from 25 historical rainfall events that triggered landslides. We compared the variation of pressure head and factor of safety with the landslide occurrence to identify the best fitting input conditions. Using calibrated inputs and a soil depth model, we ran TRIGRS for the study area. Receiver operating characteristic (ROC) analysis, comparing the model’s output with a shallow landslide inventory, shows that TRIGRS effectively simulated the instability conditions in the post-orogenic complex during historical rainfall scenarios. The implication of this work is that rainfall-induced landslides over large regions may be predicted by a deterministic model, even where data on geotechnical and hydraulic properties as well as temporal changes in topography or subsurface conditions are not available.

Esino river basin, Marche region

Estimation of natural historical flows for the Manitowish River near Manitowish Waters, Wisconsin

The Wisconsin Department of Natural Resources is charged with oversight of dam operations throughout Wisconsin and is considering modifications to the operating orders for the Rest Lake Dam in Vilas County, Wisconsin. State law requires that the operation orders be tied to natural low flows at the dam. Because the presence of the dam confounds measurement of natural flows, the U.S. Geological Survey, in cooperation with the Wisconsin Department of Natural Resources, installed streamflow-gaging stations and developed two statistical methods to improve estimates of natural flows at the Rest Lake Dam. Two independent methods were used to estimate daily natural flow for the Manitowish River approximately 1 mile downstream of the Rest Lake Dam. The first method was an adjusted drainage-area ratio method, which used a regression analysis that related measured water yield (flow divided by watershed area) from short-term (2009–11) gaging stations upstream of the Manitowish Chain of Lakes to the water yield from two nearby long-term gaging stations in order to extend the flow record (1991–2011). In this approach, the computed flows into the Chain of Lakes at the upstream gaging stations were multiplied by a coefficient to account for the monthly hydrologic contributions (precipitation, evaporation, groundwater, and runoff) associated with the additional watershed area between the upstream gaging stations and the dam at the outlet of the Chain of Lakes (Rest Lake Dam). The second method used to estimate daily natural flow at the Rest Lake Dam was a water-budget approach, which used lake stage and dam outflow data provided by the dam operator. A water-budget model was constructed and then calibrated with an automated parameter-estimation program by matching simulated flow-duration statistics with measured flow-duration statistics at the upstream gaging stations. After calibration of the water-budget model, the model was used to compute natural flow at the dam from 1973 to 2011. Daily natural flows at the dam, as computed by the adjusted drainage-area ratio method and the water-budget method, were used to compute monthly flow-duration values for the period of historical data available for each method. Monthly flow-durations provide a means for evaluating the frequency and range in flows that have been observed for each month over the course of many years. Both methods described the pattern and timing of measured high-flow and low-flow events at the upstream gaging stations. The adjusted drainage-area ratio method generally had smaller residual errors across the full range of observed flows and had smaller monthly biases than the water-budget method. Although it is not possible to evaluate which method may be more "correct" for estimating monthly natural flows at the dam, comparisons between the results of each method indicate that the adjusted drainage-area ratio method may be susceptible to biases at high flows due to isolated storms outside of the Manitowish River watershed. Conversely, it appears that the water-budget method may be susceptible to biases at low flows because of its sensitivity to the accuracy of reported lake stage and outflows, as well as effects of upstream diversions that could not be fully compensated for with this method. Results from both methods are useful for understanding the natural flow patterns at the dam. Flows for both methods have similar patterns, with high median flows in spring and low median flows in late summer. Similarly, the range from monthly high-flow durations to low-flow durations increases during spring, decreases during summer, and increases again during fall. These seasonal patterns illustrate a challenge with interpreting a single value of natural low flow. That is, a natural low flow computed for September is not representative of a natural low flow in April. Moreover, alteration of natural flows caused by storing water in the Chain of Lakes during spring and releasing it in fall causes a change in the timing of high and low flows compared with natural conditions. That is, the lowest reported dam outflows occurred in spring and highest reported outflows occurred in fall, which is opposite the natural patterns.

Wisconsin

Imputation approaches for animal movement modeling

The analysis of telemetry data is common in animal ecological studies. While the collection of telemetry data for individual animals has improved dramatically, the methods to properly account for inherent uncertainties (e.g., measurement error, dependence, barriers to movement) have lagged behind. Still, many new statistical approaches have been developed to infer unknown quantities affecting animal movement or predict movement based on telemetry data. Hierarchical statistical models are useful to account for some of the aforementioned uncertainties, as well as provide population-level inference, but they often come with an increased computational burden. For certain types of statistical models, it is straightforward to provide inference if the latent true animal trajectory is known, but challenging otherwise. In these cases, approaches related to multiple imputation have been employed to account for the uncertainty associated with our knowledge of the latent trajectory. Despite the increasing use of imputation approaches for modeling animal movement, the general sensitivity and accuracy of these methods have not been explored in detail. We provide an introduction to animal movement modeling and describe how imputation approaches may be helpful for certain types of models. We also assess the performance of imputation approaches in two simulation studies. Our simulation studies suggests that inference for model parameters directly related to the location of an individual may be more accurate than inference for parameters associated with higher-order processes such as velocity or acceleration. Finally, we apply these methods to analyze a telemetry data set involving northern fur seals ( Callorhinus ursinus ) in the Bering Sea. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme

Estimating flood magnitude and frequency at gaged and ungaged sites on streams in Alaska and conterminous basins in Canada, based on data through water year 2012

Estimates of the magnitude and frequency of floods are needed across Alaska for engineering design of transportation and water-conveyance structures, flood-insurance studies, flood-plain management, and other water-resource purposes. This report updates methods for estimating flood magnitude and frequency in Alaska and conterminous basins in Canada. Annual peak-flow data through water year 2012 were compiled from 387 streamgages on unregulated streams with at least 10 years of record. Flood-frequency estimates were computed for each streamgage using the Expected Moments Algorithm to fit a Pearson Type III distribution to the logarithms of annual peak flows. A multiple Grubbs-Beck test was used to identify potentially influential low floods in the time series of peak flows for censoring in the flood frequency analysis. For two new regional skew areas, flood-frequency estimates using station skew were computed for stations with at least 25 years of record for use in a Bayesian least-squares regression analysis to determine a regional skew value. The consideration of basin characteristics as explanatory variables for regional skew resulted in improvements in precision too small to warrant the additional model complexity, and a constant model was adopted. Regional Skew Area 1 in eastern-central Alaska had a regional skew of 0.54 and an average variance of prediction of 0.45, corresponding to an effective record length of 22 years. Regional Skew Area 2, encompassing coastal areas bordering the Gulf of Alaska, had a regional skew of 0.18 and an average variance of prediction of 0.12, corresponding to an effective record length of 59 years. Station flood-frequency estimates for study sites in regional skew areas were then recomputed using a weighted skew incorporating the station skew and regional skew. In a new regional skew exclusion area outside the regional skew areas, the density of long-record streamgages was too sparse for regional analysis and station skew was used for all estimates. Final station flood frequency estimates for all study streamgages are presented for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities. Regional multiple-regression analysis was used to produce equations for estimating flood frequency statistics from explanatory basin characteristics. Basin characteristics, including physical and climatic variables, were updated for all study streamgages using a geographical information system and geospatial source data. Screening for similar-sized nested basins eliminated hydrologically redundant sites, and screening for eligibility for analysis of explanatory variables eliminated regulated peaks, outburst peaks, and sites with indeterminate basin characteristics. An ordinary least‑squares regression used flood-frequency statistics and basin characteristics for 341 streamgages (284 in Alaska and 57 in Canada) to determine the most suitable combination of basin characteristics for a flood-frequency regression model and to explore regional grouping of streamgages for explaining variability in flood-frequency statistics across the study area. The most suitable model for explaining flood frequency used drainage area and mean annual precipitation as explanatory variables for the entire study area as a region. Final regression equations for estimating the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probability discharge in Alaska and conterminous basins in Canada were developed using a generalized least-squares regression. The average standard error of prediction for the regression equations for the various annual exceedance probabilities ranged from 69 to 82 percent, and the pseudo-coefficient of determination (pseudo-R 2 ) ranged from 85 to 91 percent. The regional regression equations from this study were incorporated into the U.S. Geological Survey StreamStats program for a limited area of the State—the Cook Inlet Basin. StreamStats is a national web-based geographic information system application that facilitates retrieval of streamflow statistics and associated information. StreamStats retrieves published data for gaged sites and, for user-selected ungaged sites, delineates drainage areas from topographic and hydrographic data, computes basin characteristics, and computes flood frequency estimates using the regional regression equations.

Alaska

A stream-gaging network analysis for the 7-day, 10-year annual low flow in New Hampshire streams

The 7-day, 10-year (7Q10) low-flow-frequency statistic is a widely used measure of surface-water availability in New Hampshire. Regression equations and basin-characteristic digital data sets were developed to help water-resource managers determine surface-water resources during periods of low flow in New Hampshire streams. These regression equations and data sets were developed to estimate streamflow statistics for the annual and seasonal low-flow-frequency, and period-of-record and seasonal period-of-record flow durations. generalized-least-squares (GLS) regression methods were used to develop the annual 7Q10 low-flow-frequency regression equation from 60 continuous-record stream-gaging stations in New Hampshire and in neighboring States. In the regression equation, the dependent variables were the annual 7Q10 flows at the 60 stream-gaging stations. The independent (or predictor) variables were objectively selected characteristics of the drainage basins that contribute flow to those stations. In contrast to ordinary-least-squares (OLS) regression analysis, GLS-developed estimating equations account for differences in length of record and spatial correlations among the flow-frequency statistics at the various stations. A total of 93 measurable drainage-basin characteristics were candidate independent variables. On the basis of several statistical parameters that were used to evaluate which combination of basin characteristics contribute the most to the predictive power of the equations, three drainage-basin characteristics were determined to be statistically significant predictors of the annual 7Q10: (1) total drainage area, (2) mean summer stream-gaging station precipitation from 1961 to 90, and (3) average mean annual basinwide temperature from 1961 to 1990. To evaluate the effectiveness of the stream-gaging network in providing regional streamflow data for the annual 7Q10, the computer program GLSNET (generalized-least-squares NETwork) was used to analyze the network by application of GLS regression between streamflow and the climatic and basin characteristics of the drainage basin upstream from each stream-gaging station. Improvement to the predictive ability of the regression equations developed for the network analyses is measured by the reduction in the average sampling-error variance, and can be achieved by collecting additional streamflow data at existing stations. The predictive ability of the regression equations is enhanced even further with the addition of new stations to the network. Continued data collection at unregulated stream-gaging stations with less than 14 years of record resulted in the greatest cost-weighted reduction to the average sampling-error variance of the annual 7Q10 regional regression equation. The addition of new stations in basins with underrepresented values for the independent variables of the total drainage area, average mean annual basinwide temperature, or mean summer stream-gaging station precipitation in the annual 7Q10 regression equation yielded a much greater cost-weighted reduction to the average sampling-error variance than when more data were collected at existing unregulated stations. To maximize the regional information obtained from the stream-gaging network for the annual 7Q10, ranking of the streamflow data can be used to determine whether an active station should be continued or if a new or discontinued station should be activated for streamflow data collection. Thus, this network analysis can help determine the costs and benefits of continuing the operation of a particular station or activating a new station at another location to predict the 7Q10 at ungaged stream reaches. The decision to discontinue an existing station or activate a new station, however, must also consider its contribution to other water-resource analyses such as flood management, water quality, or trends in land use or climatic change.

Water-Resources Investigations Report

A global ensemble of ocean wave climate statistics from contemporary wave reanalysis and hindcasts

There are numerous global ocean wave reanalysis and hindcast products currently being distributed and used across different scientific fields. However, there is not a consistent dataset that can sample across all existing products based on a standardized framework. Here, we present and describe the first coordinated multi-product ensemble of present-day global wave fields available to date. This dataset, produced through the Coordinated Ocean Wave Climate Project (COWCLIP) phase 2, includes general and extreme statistics of significant wave height ( H s ), mean wave period ( T m ) and mean wave direction ( θ m ) computed across 1980–2014, at different frequency resolutions (monthly, seasonally, and annually). This coordinated global ensemble has been derived from fourteen state-of-the-science global wave products obtained from different atmospheric reanalysis forcing and downscaling methods. This data set has been processed, under a specific framework for consistency and quality, following standard Data Reference Syntax, Directory Structures and Metadata specifications. This new comprehensive dataset provides support to future broad-scale analysis of historical wave climatology and variability as well as coastal risk and vulnerability assessments across offshore and coastal engineering applications.

Scientific Data

Principles of logic and the use of digital geographic information systems

Digital geographic information systems allow many different types of data to be spatially and statistically analyzed. Logical operations can be performed on individual or multiple data planes by algorithms that can be implemented in computer systems. Users and creators of the systems should fully understand these operations. This paper describes the relationships of layers and features in geographic data bases and the principles of logic that can be applied by geographic information systems and suggests that a thorough knowledge of the data that are entered into a geographic data base and of the logical operations will produce results that are most satisfactory to the user. Methods of spatial analysis are reduced to their primitive logical operations and explained to further such understanding.

Circular

Areal distribution and concentration of contaminants of concern in surficial streambed and lakebed sediments, Lake St. Clair and tributaries, Michigan, 1990-2003

As part of the Lake St. Clair Regional Monitoring Project, the U.S. Geological Survey evaluated data collected from surficial streambed and lakebed sediments in the Lake Erie-Lake St. Clair drainages. This study incorporates data collected from 1990 through 2003 and focuses primarily on the U.S. part of the Lake St. Clair Basin, including Lake St. Clair, the St. Clair River, and tributaries to Lake St. Clair. Comparable data from the Canadian part of the study area are included where available. The data are compiled into 4 chemical classes and consist of 21 compounds. The data are compared to effects-based sediment-quality guidelines, where the Threshold Effect Level and Lowest Effect Level represent concentrations below which adverse effects on biota are not expected and the Probable Effect Level and Severe Effect Level represent concentrations above which adverse effects on biota are expected to be frequent. Maps in the report show the spatial distribution of the sampling locations and illustrate the concentrations relative to the selected sediment-quality guidelines. These maps indicate that sediment samples from certain areas routinely had contaminant concentrations greater than the Threshold Effect Concentration or Lowest Effect Level. These locations are the upper reach of the St. Clair River, the main stem and mouth of the Clinton River, Big Beaver Creek, Red Run, and Paint Creek. Maps also indicated areas that routinely contained sediment contaminant concentrations that were greater than the Probable Effect Concentration or Severe Effect Level. These locations include the upper reach of the St. Clair River, the main stem and mouth of the Clinton River, Red Run, within direct tributaries along Lake St. Clair and in marinas within the lake, and within the Clinton River headwaters in Oakland County. Although most samples collected within Lake St. Clair were from sites adjacent to the mouths of its tributaries, samples analyzed for trace-element concentrations were collected throughout the lake. The distribution of trace-element concentrations corresponded well with the results of a two-dimensional hydrodynamic model of flow patterns from the Clinton River into Lake St. Clair. The model was developed independent from the bed sediment analysis described in this report; yet it showed a zone of deposition for outflow from the Clinton River into Lake St. Clair that corresponded well with the spatial distribution of trace-element concentrations. This zone runs along the western shoreline of Lake St. Clair from L'Anse Creuse Bay to St. Clair Shores, Michigan and is reflected in the samples analyzed for mercury and cadmium. Statistical summaries of the concentration data are presented for most contaminants, and selected statistics are compared to effects-based sediment-quality guidelines. Summaries were not computed for dieldrin, chlordane, hexachlorocyclohexane, lindane, and mirex because insufficient data are available for these contaminants. A statistical comparison showed that the median concentration for hexachlorobenzene, anthracene, benz[ a ]anthracene, chrysene, and pyrene are greater than the Threshold Effect Concentration or Lowest Effect Level. Probable Effect Concentration Quotients provide a mechanism for comparing the concentrations of contaminant mixtures against effects-based biota data. Probable Effect Concentration Quotients were calculated for individual samples and compared to effects-based toxicity ranges. The toxicity-range categories used in this study were nontoxic (quotients < 0.5) and toxic (quotients > 0.5). Of the 546 individual samples for which Probable Effect Concentration Quotients were calculated, 469 (86 percent) were categorized as being nontoxic and 77 (14 percent) were categorized as being toxic. Bed-sediment samples with toxic Probable Effect Concentration Quotients were collected from Paint Creek, Galloway Creek, the main stem of the Clinton River, Big Beaver Creek, Red Run, Clinton River towards the mouth, Lake St. Clair along the western shore, and the St. Clair River near Sarnia.

Michigan

Reference manual for generation and analysis of Habitat Time Series: version II

The selection of an instream flow requirement for water resource management often requires the review of how the physical habitat changes through time. This review is referred to as 'Time Series Analysis." The Tune Series Library (fSLIB) is a group of programs to enter, transform, analyze, and display time series data for use in stream habitat assessment. A time series may be defined as a sequence of data recorded or calculated over time. Examples might be historical monthly flow, predicted monthly weighted usable area, daily electrical power generation, annual irrigation diversion, and so forth. The time series can be analyzed, both descriptively and analytically, to understand the importance of the variation in the events over time. This is especially useful in the development of instream flow needs based on habitat availability. The TSLIB group of programs assumes that you have an adequate study plan to guide you in your analysis. You need to already have knowledge about such things as time period and time step, species and life stages to consider, and appropriate comparisons or statistics to be produced and displayed or tabulated. Knowing your destination, you must first evaluate whether TSLIB can get you there. Remember, data are not answers . This publication is a reference manual to TSLIB and is intended to be a guide to the process of using the various programs in TSLIB. This manual is essentially limited to the hands-on use of the various programs. a TSLIB use interface program (called RTSM) has been developed to provide an integrated working environment where the use has a brief on-line description of each TSLIB program with the capability to run the TSLIB program while in the user interface. For information on the RTSM program, refer to Appendix F. Before applying the computer models described herein, it is recommended that the user enroll in the short course "Problem Solving with the Instream Flow Incremental Methodology (IFIM)." This course is offered by the Aquatic Systems Branch of the National Ecology Research Center. For more information about the TSLIB software, refer to the Memorandum of Understanding. Chapter 1 provides a brief introduction to the Instream Flow Incremental Methodology and TSLIB. Other chapters in this manual provide information on the different aspects of using the models. The information contained in the other chapters includes (2) acquisition, entry, manipulation, and listing of streamflow data; (3) entry, manipulation, and listing of the habitat-versus-streamflow function; (4) transferring streamflow data; (5) water resources systems analysis; (6) generation and analysis of daily streamflow and habitat values; (7) generation of the time series of monthly habitats; (8) manipulation, analysis, and display of month time series data; and (9) generation, analysis, and display of annual time series data. Each section includes documentation for the programs therein with at least one page of information for each program, including a program description, instructions for running the program, and sample output. The Appendixes contain the following: (A) sample file formats; (B) descriptions of default filenames; (C) alphabetical summary of batch-procedure files; (D) installing and running TSLIB on a microcomputer; (E) running TSLIB on a CDC Cyber computer; (F) using the TSLIB user interface program (RTSM); and (G) running WATSTORE on the USGS Amdahl mainframe computer. The number for this version of TSLIB--Version II-- is somewhat arbitrary, as the TSLIB programs were collected into a library some time ago; but operators tended to use and manage them as individual programs. Therefore, we will consider the group of programs from the past that were only on the CDC Cyber computer as Version 0; the programs from the past that were on both the Cyber and the IBM-compatible microcomputer as Version I; and the programs contained in this reference manual as Version II.

Report

Application of a workflow to determine the feasibility of using simulated streamflow for estimation of streamflow frequency statistics

Streamflow records from hydrologic models are attractive for use in operational hydrology, such as a streamflow frequency analysis. The amount of bias inherent to simulated streamflow from hydrologic models is often unknown, but it is likely present in derivative products. Therefore, a workflow may help determine where streamflow frequency analysis is credibly feasible from simulated streamflow and allow for a systematic way to assess and correct for bias. The proposed workflow consists of hydrologically matching model output locations with streamflow-gauging station (stream gauge) locations, computing the desired statistic from the simulated and observed streamflow record, computing the differences between the simulated and observed statistic (i.e., the bias), and constructing generalized additive models (GAMs) from the differences to determine bias corrections. The US Geological Survey, in cooperation with the Gulf Coast Ecosystem Restoration Council and the US Environmental Protection Agency, is testing the proposed workflow on a low-streamflow frequency (LFF) analysis. Simulated streamflows for the LFF analysis were sourced from a machine-learning model that estimated daily streamflow at Level-12 hydrologic unit code (HUC12) pour points (outlets) in the southern and southeastern US for 1950–2010. The comparison data set consists of 497 stream gauges that are coincident with a HUC12 outlet. The simulated LFF statistics were being overestimated on average; thus, there are limits to using simulated streamflow for frequency analysis. The magnitude of the overprediction generally increases where no-flow conditions are common. Bias corrections determined from the GAMs decreased the magnitude of bias observed in the simulated LFF statistics on average, suggesting it is feasible to expand the operational use of simulated streamflows to frequency analyses with the proposed workflow. The proposed workflow could be advantageous to practitioners interested in leveraging existing and future simulated streamflow data sets with regional and or global coverage.

Journal of Hydrologic Engineering

Hydrologic budget of the Beaverdam Creek basin, Maryland

A hydrologic budget is a statement accounting for the water gains and losses for selected periods in an area. Weekly measurements of precipitation streamflow, surface-water storage, ground-water stage, and soil resistivity were made during a 2year period, April 1, 1950, to March 28, 1952, in the Beaverdam Creek basin, Wicomico County, Md. The hydrologic measurements are summarized in two budgets, a total budget and a ground-water budget, and in supporting tables and graphs. The results of the investigation have some potentially significant applications because they describe a method for determining the annual replenishment of the water supply of a basin and the ways of water disposal under natural conditions. The information helps to determine the 'safe' yield of water in diversion from natural to artificial discharge. The drainage basin of Beaverdam Creek was selected because it appeared to have fewer hydrologic variables than are generally found. However, the methods may prove applicable in many places under a variety of conditions. The measurements are expressed in inches of water over the area of the basin. The equation of the hydrologic cycle is the budget balance: P= R+E+ASW+ delta SW + delta SM + delta GW where P is precipitation; R is runoff; ET is evapotranspiration; delta SW is change in surface-water storage; delta SM is change in soil moisture; and delta GW is change in ground-water storage. In this report 'change' is the final quantity minus the initial quantity and thus is synonymous with 'increase.' Further, ,delta GW= delta H .x Yg, in which delta H is the change in ground-water stage and Yg is the gravity yield, or the specific yield of the sediments as measured during the short periods of declining ground-water levels characteristic of the area. The complex sum of the revised equation P ? R - delta SW ? ET - delta SM, which is equal to delta H. x Yg, has been named the 'infiltration residual'; it is equivalent to ground-water recharge. Two unmeasured, but not entirely unknown, quantities, evapotranspiration, (ET) and gravity yield, (Yg), are included in the equation. They are derived statistically by a method of convergent approximations, one of the contributions of this investigation. On the basis of laboratory analysis, well-field tests, and general information on rates of drainage from saturated sediments, a gravity yield of 14 percent was assumed as a first approximation. The equation was then solved, by weeks, for evapotranspiration, ET. The evapotranspiration losses were plotted against the calendar week. Using the time of year as a control, a smooth curve was fitted to the evapotranspiration data, and modified values of ET were read from the curve. These were used to compute weekly values of the infiltration residual which were plotted against ground-water stage. The slope of the line of best fit gave a closer approximation of gravity yield, Yg. The process was repeated. The approximations converged, so that a fourth and final approximation resulted in a close grouping of all the points along a line whose slope indicated a Yg of 11.0 percent, and a slightly asymmetric bell-shaped curve of total evapotranspiration by weeks was obtained that is considered representative of this area. Check calculations of gravity yield were made during periods of low evapotranspiration and high infiltration, which substantiate the computed average of 11.0 percent. Refinements in the method of deriving the ground-water budget were introduced to supplement the techniques developed by Meinzer and Stearns in the study of the Pomperaug River basin in Connecticut in 1913 and 1916. The hydrologic equation for the ground-water cycle may be written Gr=D + delta H. x Yg + ETg, in which Gr is ground-water recharge (infiltration); D is ground-water drainage; delta H is the change in mean ground-water stage (final stage minus initial stage); Yg is gravity yield (taken as 11.0 percent in computations here); an

Water Supply Paper