Geology ReportsSearch

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Considerations for blending data from various sensors

A project is being proposed at the EROS Data Center to blend the information from sensors aboard various satellites. The problems of, and considerations for, blending data from several satellite-borne sensors are discussed. System descriptions of the sensors aboard the HCMM, TIROS-N, GOES-D, Landsat 3, Landsat D, Seasat, SPOT, Stereosat, and NOSS satellites, and the quantity, quality, image dimensions, and availability of these data are summaries to define attributes of a multi-sensor satellite data base. Unique configurations of equipment, storage, media, and specialized hardware to meet the data system requirement are described as well as archival media and improved sensors that will be on-line within the next 5 years. Definitions and rigor required for blending various sensor data are given. Problems of merging data from the same sensor (intrasensor comparison) and from different sensors (intersensor comparison), the characteristics and advantages of cross-calibration of data, and integration of data into a product matrix field are addressed. Data processing considerations as affected by formation, resolution, and problems of merging large data sets, and organization of data bases for blending data are presented. Examples utilizing GOES and Landsat data are presented to demonstrate techniques of data blending, and recommendations for future implementation of a set of standard scenes and their characteristics necessary for optimal data blending are discussed.

Sixth Annual Pecora Symposium and Exposition

Alaska Geochemical Database, Version 2.0 (AGDB2)–Including “best value” data compilations for rock, sediment, soil, mineral, and concentrate sample mediaI

The Alaska Geochemical Database Version 2.0 (AGDB2) contains new geochemical data compilations in which each geologic material sample has one “best value” determination for each analyzed species, greatly improving speed and efficiency of use. Like the Alaska Geochemical Database (AGDB, http://pubs.usgs.gov/ds/637/) before it, the AGDB2 was created and designed to compile and integrate geochemical data from Alaska in order to facilitate geologic mapping, petrologic studies, mineral resource assessments, definition of geochemical baseline values and statistics, environmental impact assessments, and studies in medical geology. This relational database, created from the Alaska Geochemical Database (AGDB) that was released in 2011, serves as a data archive in support of present and future Alaskan geologic and geochemical projects, and contains data tables in several different formats describing historical and new quantitative and qualitative geochemical analyses. The analytical results were determined by 85 laboratory and field analytical methods on 264,095 rock, sediment, soil, mineral and heavy-mineral concentrate samples. Most samples were collected by U.S. Geological Survey personnel and analyzed in U.S. Geological Survey laboratories or, under contracts, in commercial analytical laboratories. These data represent analyses of samples collected as part of various U.S. Geological Survey programs and projects from 1962 through 2009. In addition, mineralogical data from 18,138 nonmagnetic heavy-mineral concentrate samples are included in this database. The AGDB2 includes historical geochemical data originally archived in the U.S. Geological Survey Rock Analysis Storage System (RASS) database, used from the mid-1960s through the late 1980s and the U.S. Geological Survey PLUTO database used from the mid-1970s through the mid-1990s. All of these data are currently maintained in the National Geochemical Database (NGDB). Retrievals from the NGDB were used to generate most of the AGDB data set. These data were checked for accuracy regarding sample location, sample media type, and analytical methods used. This arduous process of reviewing, verifying and, where necessary, editing all U.S. Geological Survey geochemical data resulted in a significantly improved Alaska geochemical dataset. USGS data that were not previously in the NGDB because the data predate the earliest U.S. Geological Survey geochemical databases, or were once excluded for programmatic reasons, are included here in the AGDB2 and will be added to the NGDB. The AGDB2 data provided here are the most accurate and complete to date, and should be useful for a wide variety of geochemical studies. The AGDB2 data provided in the linked database may be updated or changed periodically.

Alaska

Water resources data, New York, water year 1996; Volume 1. Eastern New York; excluding Long Island

Introduction Water-resources data for the 1996 water year for New York consist of records of stage, discharge, and water quality of streams; stage, contents, and water quality of lakes and reservoirs; ground-water levels; and precipitation quality. This volume contains records for water discharge at 122 gaging stations; stage only at 7 gaging stations; stage and contents at 4 gaging stations, and 18 other lakes and reservoirs; water quality at 28 gaging stations and 1 precipitation-quality station; and water levels at 3 observation wells. Also included are data for 33 crest-stage partial-record stations. Additional water data were collected at various sites not involved in the systematic data-collection program, and are published as miscellaneous measurements and analyses in this volume. These data together with the data in Volumes 2 and 3 represent that part of the National Water Data System operated by the U.S. Geological Survey in cooperation with State, Municipal, and Federal agencies in New York. Records of discharge and stage of streams, and contents and stage of lakes and reservoirs, were first published in a series of U.S. Geological Survey water-supply papers entitled, “Surface Water Supply of the United States.” Through September 30, 1960, these water-supply papers were in an annual series and then in a 5-year series for 1961-65 and 1966-70. Records of water quality, water temperatures, and suspended sediment were published from 1941 to 1970 in an annual series of water-supply papers entitled “Quality of Surface Waters of the United States.” Records of ground-water levels were published from 1935 to 1974 in a series of water-supply papers entitled “Ground-Water Levels in the United States.” Water-supply papers may be consulted in the libraries of the principal cities and universities in the United States or may be purchased from the U.S. Geological Survey, Branch of Distribution, 604 South Pickett Street, Alexandria, VA 22304. Since the 1961 water year, streamflow data and since the 1964 water year, water-quality data have been released by the Geological Survey in annual reports on a State-boundary basis. These reports provided rapid release of water data in each state shortly after the end of the water year. Through 1970 the data were also released in the water-supply paper series mentioned above. Streamflow and water-quality data beginning with the 1971 water year, and ground-water data beginning with the 1975 water year are published only in reports on a State-boundary basis. Beginning with the 1975 water year, these Survey reports carry an identification number consisting of the two-letter State abbreviation, the last two digits of the water year, and the volume number. For example, this volume is identified as “U.S. Geological Survey Water-Data Report NY-96-1.” Water-data reports are for sale in paper copy or in microfiche by the National Technical Information Service, U.S. Department of Commerce, Springfield, VA 22161. Additional information, including current prices for ordering specific reports, may be obtained from the District Office at the address given on the back of the title page or by telephone (518) 285-5600.

New York

Estimating the drivers of species distributions with opportunistic data using mediation analysis

Ecological occupancy modeling has historically relied on high-quality, low-quantity designed-survey data for estimation and prediction. In recent years, there has been a large increase in the amount of high-quantity, unknown-quality opportunistic data. This has motivated research on how best to combine these two data sources in order to optimize inference. Existing methods can be infeasible for large datasets or require opportunistic data to be located where designed-survey data exist. These methods map species occupancies, motivating a need to properly evaluate covariate effects (e.g., land cover proportion) on their distributions. We describe a spatial estimation method for supplementarily including additional opportunistic data using mediation analysis concepts. The opportunistic data mediate the effect of the covariate on the designed-survey data response, decomposing it into a direct and indirect effect. A component of the indirect effect can then be quickly estimated via regressing the mediator on the covariate, while the other components are estimated through a spatial occupancy model. The regression step allows for use of large quantities of opportunistic data that can be collected in locations with no designed-survey data available. Simulation results suggest that the mediated method produces an improvement in relative MSE when the data are of reasonable quality. However, when the simulated opportunistic data are poorly correlated with the true spatial process, the standard, unmediated method is still preferable. A spatiotemporal extension of the method is also developed for analyzing the effect of deciduous forest land cover on red-eyed vireo distribution in the southeastern United States and find that including the opportunistic data do not lead to a substantial improvement. Opportunistic data quality remains an important consideration when employing this method, as with other data integration methods.

eastern United States

Spatial data reduction through element -of-interest (EOI) extraction

Any large, multifaceted data collection that is challenging to handle with traditional management practices can be branded ‘Big Data.’ Any big data containing geo-referenced attributes can be considered big geospatial data. The increased proliferation of big geospatial data is currently reforming the geospatial industry into a data-driven enterprise. Challenges in the big spatial data domain can be summarized as the ‘Big Vs’ – variety, volume, velocity, veracity and value. Big spatial data sources can be considered in two broad classes, active and passive, as each is impacted to varying degrees. Some of these challenges may be alleviated by reducing unprocessed, or minimally processed, (raw) data to features, which we refer to as the extraction of Elements of Interest (EOI). In fact, many applications require EOI extraction from raw data to enable their basic employment. This chapter presents current state-of-the-art methods to create EOI from some types of georeferenced big data. We classify the data types into two realms: active and passive. Active data are those collected specifically for the purpose to which they are applied. Passive data are those collected for purposes other than those for which they are utilized, included those ‘collected’ for no particular purpose at all. The chapter then presents use cases from both the active and passive spatial realms, including the active applications of terrain feature extraction from digital elevation models and vegetation mapping from remotely-sensed imagery and passive applications like building identification from VGI and point-of-interest data mining from social networks for land use classification. Finally, the chapter concludes with future research needs.

Book chapter

Data standardization and management to facilitate large-scale and interdisciplinary approaches access

Bringing data related to recreational fishers and fisheries together across large scales can provide tremendous insight. Methods for collecting, analysing, and storing data can vary dramatically, which can have significant implications for the use of these data. Efforts to standardise data within organisations often increase the ability to compare datasets from different areas, monitor changes over time, and increase the utility of the data for management and research. Though employing standardised methodology and data architecture results in the most straightforward and robust opportunities for data integration, doing so may not be possible due to variation in data collection objectives, continuity with historical programs, and resource limitations. Additionally, managing data according to FAIR principles (findability, accessibility, interoperability, and reusability) helps support data sharing and large-scale research efforts. Key elements of integrable data include appropriate data structures, adequate documentation of methodology, interpretable and complete metadata, and accessible storage formats. We document the potential benefits of standardising data and offer example approaches. We also explore best practices regarding the formatting, storage, and transmission of recreational fisher data. Efforts to increase the level of standardisation and integrability of recreational fisher data can create opportunities to better understand fisher behaviour, needs, and fulfilment.

Book chapter

Sharing FAIR monitoring program data improves discoverability and reuse

Data resulting from environmental monitoring programs are valuable assets for natural resource managers, decision-makers, and researchers. These data are often collected to inform specific reporting needs or decisions with a specific timeframe. While program-oriented data and related publications are effective for meeting program goals, sharing well-documented data and metadata allows users to research aspects outside initial program intentions. As part of an effort to integrate data from four long-term large-scale US aquatic monitoring programs, we evaluated the original datasets against the FAIR (Findable, Accessible, Interoperable, Reusable) data principles and offer recommendations and lessons learned. Differences in data governance across these programs resulted in considerable effort to access and reuse the original datasets. Requirements, guidance, and resources available to support data publishing and documentation are inconsistent across agencies and monitoring programs, resulting in various data formats and storage locations that are not easily found, accessed, or reused. Making monitoring data FAIR will reduce barriers to data discovery and reuse. Programs are continuously striving to improve data management, data products, and metadata; however, provision of related tools, consistent guidelines and standards, and more resources to do this work is needed. Given the value of these data and the significant effort required to access and reuse them, actions and steps intended on improving data documentation and accessibility are described.

Enviornmental Monitoring and Assessment

Organization of marine phenology data in support of planning and conservation in ocean and coastal ecosystems

Among the many effects of climate change is its influence on the phenology of biota. In marine and coastal ecosystems, phenological shifts have been documented for multiple life forms; however, biological data related to marine species' phenology remain difficult to access and is under-used. We conducted an assessment of potential sources of biological data for marine species and their availability for use in phenological analyses and assessments. Our evaluations showed that data potentially related to understanding marine species' phenology are available through online resources of governmental, academic, and non-governmental organizations, but appropriate datasets are often difficult to discover and access, presenting opportunities for scientific infrastructure improvement. The developing Federal Marine Data Architecture when fully implemented will improve data flow and standardization for marine data within major federal repositories and provide an archival repository for collaborating academic and public data contributors. Another opportunity, largely untapped, is the engagement of citizen scientists in standardized collection of marine phenology data and contribution of these data to established data flows. Use of metadata with marine phenology related keywords could improve discovery and access to appropriate datasets. When data originators choose to self-publish, publication of research datasets with a digital object identifier, linked to metadata, will also improve subsequent discovery and access. Phenological changes in the marine environment will affect human economics, food systems, and recreation. No one source of data will be sufficient to understand these changes. The collective attention of marine data collectors is needed—whether with an agency, an educational institution, or a citizen scientist group—toward adopting the data management processes and standards needed to ensure availability of sufficient and useable marine data to understand marine phenology.

Ecological Informatics

Challenges with secondary use of multi-source water-quality data in the United States

Combining water-quality data from multiple sources can help counterbalance diminishing resources for stream monitoring in the United States and lead to important regional and national insights that would not otherwise be possible. Individual monitoring organizations understand their own data very well, but issues can arise when their data are combined with data from other organizations that have used different methods for reporting the same common metadata elements. Such use of multi-source data is termed “secondary use”—the use of data beyond the original intent determined by the organization that collected the data. In this study, we surveyed more than 25 million nutrient records collected by 488 organizations in the United States since 1899 to identify major inconsistencies in metadata elements that limit the secondary use of multi-source data. Nearly 14.5 million of these records had missing or ambiguous information for one or more key metadata elements, including (in decreasing order of records affected) sample fraction, chemical form, parameter name, units of measurement, precise numerical value, and remark codes. As a result, metadata harmonization to make secondary use of these multi-source data will be time consuming, expensive, and inexact. Different data users may make different assumptions about the same ambiguous data, potentially resulting in different conclusions about important environmental issues. The value of these ambiguous data is estimated at \$US12 billion, a substantial collective investment by water-resource organizations in the United States. By comparison, the value of unambiguous data is estimated at \$US8.2 billion. The ambiguous data could be preserved for uses beyond the original intent by developing and implementing standardized metadata practices for future and legacy water-quality data throughout the United States.

Water Research

Stakeholder engagement to guide decision-relevant water data delivery

Water resources management and policy making require access to reliable scientific data. However, water managers may need to overcome various obstacles to accessing data. For example, insufficient technological infrastructures, low data literacy, and data format complexities often inhibit data user access. Thus, it is imperative to include stakeholders in the design of data delivery systems. The United States Geological Survey's Water Resources Mission Area is currently developing Integrated Water Availability Assessments (IWAAs) — multi-extent, stakeholder driven, near real-time water availability census and prediction for human and ecological uses. To provide appropriate user accessibility to data delivery systems developed for IWAAs, a user-centered design process including stakeholder focus groups was used to determine potential water data user needs and preferences. Focus groups identified five types of potential users: Public sector water resources managers, Public sector water resources manager data analysts, Industry and private companies, Tribal Nations, and Nonprofit organizations. Different water data user types depended on diverse spatial and temporal scale data. Public sector water resources managers benefitted most from data synthesized into user-friendly platforms and Public sector water resources data analysts preferred easy access to raw data. These findings can support the development of a water data delivery platform that meets a variety of user needs.

Journal of the American Water Resources Associatio

What you should know about land-cover data

Wildlife biologists are using land-characteristics data sets for a variety of applications. Many kinds of landscape variables have been characterized and the resultant data sets or maps are readily accessible. Often, too little consideration is given to the accuracy or traits of these data sets, most likely because biologists do not know how such data are compiled and rendered, or the potential pitfalls that can be encountered when applying these data. To increase understanding of the nature of land-characteristics data sets, I introduce aspects of source information and data-handling methodology that include the following: ambiguity of land characteristics; temporal considerations and the dynamic nature of the landscape; type of source data versus landscape features of interest; data resolution, scale, and geographic extent; data entry and positional problems; rare landscape features; and interpreter variation. I also include guidance for determining the quality of land-characteristics data sets through metadata or published documentation, visual clues, and independent information. The quality or suitability of the data sets for wildlife applications may be improved with thematic or spatial generalization, avoidance of transitional areas on maps, and merging of multiple data sources. Knowledge of the underlying challenges in compiling such data sets will help wildlife biologists to better assess the strengths and limitations and determine how best to use these data.

Journal of Wildlife Management

Digital data sets for map products produced as part of the Black Hills Hydrology Study, western South Dakota

This compact disk contains digital data produced as part of the 1:100,000-scale map products for the Black Hills Hydrology Study conducted in western South Dakota. The digital data include 28 individual Geographic Information System (GIS) data sets: data sets for the hydrogeologic unit map including all mapped hydrogeologic units within the study area (1 data set) and major geologic structure including anticlines and synclines (1 data set); data sets for potentiometric maps including the potentiometric contours for the Inyan Kara, Minnekahta, Minnelusa, Madison, and Deadwood aquifers (5 data sets), wells used as control points for each aquifer (5 data sets), and springs used as control points for the potentiometric contours (1 data set); and data sets for the structure-contour maps including the structure contours for the top of each formation that contains major aquifers (5 data sets), wells and tests holes used as control points for each formation (5 data sets), and surficial deposits (alluvium and terrace deposits) that directly overlie each of the major aquifer outcrops (5 data sets). These data sets were used to produce the maps published by the U.S. Geological Survey.

Open-File Report

Global multi-resolution terrain elevation data 2010 (GMTED2010)

In 1996, the U.S. Geological Survey (USGS) developed a global topographic elevation model designated as GTOPO30 at a horizontal resolution of 30 arc-seconds for the entire Earth. Because no single source of topographic information covered the entire land surface, GTOPO30 was derived from eight raster and vector sources that included a substantial amount of U.S. Defense Mapping Agency data. The quality of the elevation data in GTOPO30 varies widely; there are no spatially-referenced metadata, and the major topographic features such as ridgelines and valleys are not well represented. Despite its coarse resolution and limited attributes, GTOPO30 has been widely used for a variety of hydrological, climatological, and geomorphological applications as well as military applications, where a regional, continental, or global scale topographic model is required. These applications have ranged from delineating drainage networks and watersheds to using digital elevation data for the extraction of topographic structure and three-dimensional (3D) visualization exercises (Jenson and Domingue, 1988; Verdin and Greenlee, 1996; Lehner and others, 2008). Many of the fundamental geophysical processes active at the Earth's surface are controlled or strongly influenced by topography, thus the critical need for high-quality terrain data (Gesch, 1994). U.S. Department of Defense requirements for mission planning, geographic registration of remotely sensed imagery, terrain visualization, and map production are similarly dependent on global topographic data. Since the time GTOPO30 was completed, the availability of higher-quality elevation data over large geographic areas has improved markedly. New data sources include global Digital Terrain Elevation Data (DTEDRegistered) from the Shuttle Radar Topography Mission (SRTM), Canadian elevation data, and data from the Ice, Cloud, and land Elevation Satellite (ICESat). Given the widespread use of GTOPO30 and the equivalent 30-arc-second DTEDRegistered level 0, the USGS and the National Geospatial-Intelligence Agency (NGA) have collaborated to produce an enhanced replacement for GTOPO30, the Global Land One-km Base Elevation (GLOBE) model and other comparable 30-arc-second-resolution global models, using the best available data. The new model is called the Global Multi-resolution Terrain Elevation Data 2010, or GMTED2010 for short. This suite of products at three different resolutions (approximately 1,000, 500, and 250 meters) is designed to support many applications directly by providing users with generic products (for example, maximum, minimum, and median elevations) that have been derived directly from the raw input data that would not be available to the general user or would be very costly and time-consuming to produce for individual applications. The source of all the elevation data is captured in metadata for reference purposes. It is also hoped that as better data become available in the future, the GMTED2010 model will be updated.

Open-File Report

Defining a data management strategy for USGS Chesapeake Bay studies

The mission of U.S. Geological Survey’s (USGS) Chesapeake Bay studies is to provide integrated science for improved understanding and management of the Chesapeake Bay ecosystem. Collective USGS efforts in the Chesapeake Bay watershed began in the 1980s, and by the mid-1990s the USGS adopted the watershed as one of its national place-based study areas. Great focus and effort by the USGS have been directed toward Chesapeake Bay studies for almost three decades. The USGS plays a key role in using “ecosystem-based adaptive management, which will provide science to improve the efficiency and accountability of Chesapeake Bay Program activities” (Phillips, 2011). Each year USGS Chesapeake Bay studies produce published research, monitoring data, and models addressing aspects of bay restoration such as, but not limited to, fish health, water quality, land-cover change, and habitat loss. The USGS is responsible for collaborating and sharing this information with other Federal agencies and partners as described under the President’s Executive Order 13508—Strategy for Protecting and Restoring the Chesapeake Bay Watershed signed by President Obama in 2009. Historically, the USGS Chesapeake Bay studies have relied on national USGS databases to store only major nationally available sources of data such as streamflow and water-quality data collected through local monitoring programs and projects, leaving a multitude of other important project data out of the data management process. This practice has led to inefficient methods of finding Chesapeake Bay studies data and underutilization of data resources. Data management by definition is “the business functions that develop and execute plans, policies, practices and projects that acquire, control, protect, deliver and enhance the value of data and information.” (Mosley, 2008a). In other words, data management is a way to preserve, integrate, and share data to address the needs of the Chesapeake Bay studies to better manage data resources, work more efficiently with partners, and facilitate holistic watershed science. It is now the goal of the USGS Chesapeake Bay studies to implement an enhanced and all-encompassing approach to data management. This report discusses preliminary efforts to implement a physical data management system for program data that is not replicated nationally through other USGS databases.

Chesapeake Bay

Community for Data Integration 2013 Annual Report

The U.S. Geological Survey (USGS) conducts earth science to help address complex issues affecting society and the environment. In 2006, the USGS held the first Scientific Information Management Workshop to bring together staff from across the organization to discuss the data and information management issues affecting the integration and delivery of earth science research and investigate the use of “communities of practice” as mechanisms to share expertise about these issues. Out of this effort emerged the Council for Data Integration, which was conceived as an official organizational function that would help guide data integration activities and formalize communities of practice into working groups. However by 2009, it became apparent that many members of the council had an interest in developing data integration solutions and sharing expertise in a less formal grassroots perspective, thus transforming the “Council” into a “Community” for Data Integration (CDI). Today, the CDI represents a dynamic community of practice focused on advancing science data and information management and integration capabilities across the USGS and the CDI community. The CDI fosters an environment for collaboration and sharing by bringing together expertise from external partners and representatives across USGS who are involved in research, data management, and information technology. Membership is voluntary and open to USGS employees and other individuals and organizations willing to contribute to the community (if interested, contact cdi@usgs.gov). The purpose of the CDI is to advance understanding of Earth systems through enhanced use of data and information including associated tools and techniques provide a forum for people doing work with data integration to come together to share ideas as well as learn new skills and techniques, and grow overall USGS capabilities with data and information by increasing visibility of the work of many people throughout the USGS and the CDI community. To achieve these goals, the CDI operates within four applied areas: monthly forums, annual workshop/webinar series, working groups, and projects. The monthly forums, also known as the Opportunity/Challenge of the Month, provide an open dialogue to share and learn about data integration efforts or to present problems that invite the Community to offer solutions, advice, and support. Since 2010, the CDI has also sponsored annual workshops/webinar series to encourage the exchange of ideas, sharing of activities, presentations of current projects, and networking among members. Stemming from common interests, the working groups are focused on efforts to address data management and technical 2 challenges, including the development of standards and tools, improving interoperability and information infrastructure, and data preservation within USGS and its partners. The growing support for the activities of the working groups led to the CDI’s first formal request for proposals (RFP) process in 2013 to fund projects that produced tangible products. Today the CDI continues to hold an annual RFP that create data management tools and practices, collaboration tools, and training in support of data integration and delivery.

Open-File Report

U.S. Geological Survey Community for Data Integration 2017 Workshop Proceedings

Executive Summary The U.S. Geological Survey (USGS) Community for Data Integration (CDI) Workshop was held May 16–19, 2017 at the Denver Federal Center. There were 183 in-person attendees and 35 virtual attendees over four days. The theme of the workshop was “Enabling Integrated Science,” with the purpose of bringing together the community to discuss current topics, shared challenges, and steps forward to advance integrated science at the USGS. The CDI welcomed several keynote speakers, including Bill Werkheiser, USGS Acting Director; Kevin T. Gallagher, USGS Associate Director of the Core Science Systems Mission Area; Bruce Caron, Earth Science Information Partners Community Architect; and Tim Quinn, Chief of the USGS Office of Enterprise Information. Their presentations focused on the importance of collaborative, cross-disciplinary, and open science and the role of the CDI in identifying and supporting new opportunities in these areas for the USGS and its partners. In addition to the stated theme, the workshop agenda was driven by the needs of the CDI, with topics highlighting current resources and technologies that could help attendees in their daily work. Topical sessions were proposed by CDI members and included subjects such as data citation, information technology architecture, legacy data, real-time data, and many more. Plenary speakers from the community talked about USGS activities in data science, elevation and hydrography data integration, advanced scientific computing solutions, cloud computing, data-management training, and data-sharing agreements. Two panels addressed the role of the CDI in enabling integrated science and examples of CDI-supported projects in action. Breakout discussions focused on the workshop theme of “Enabling Integrated Science” and covered five topics: Data and Data Integration, Modeling, Computing Capacity, Science Data Integration, and User Needs and Experience. Sessions on each topic identified actions that could bring the USGS and the broader Earth science community closer to the goal of making integrated science commonplace. The breakouts produced recommendations with the broad themes of improving communication and connections across the USGS, reducing duplication and increasing knowledge transfer, increasing training and testbed opportunities to learn and experiment, and creating community-supported standards to enable better integration and interoperability. The DataBlast poster and live demonstration session showcased 36 projects from around the CDI and included recent CDI-funded projects as well as other USGS and partner initiatives that were related to data and software integration and discovery. Importantly, the CDI workshop provided a forum for scientists, technologists, data and resource managers, program managers, and others to convene face to face to discuss common methods, interests, challenges, and solutions related to scientific data and technologies. As a result of this rare convergence, new connections were made across disciplines, backgrounds, and geographical locations, seeding future activities and collaborations. Sharing of ideas from all attendees was encouraged through the use of a mobile application to collect real-time questions and feedback from the audience The primary outcomes of the workshop are the recommendations from the breakout sessions titled “Roadmap Discussions on Enabling Integrated Science” and from the topical sessions detailed in these proceedings. These sessions, as well as the plenary discussions, identified new areas of collaboration and learning that the CDI will facilitate, such as data science, software development, scientific modeling practices, and user needs and experience. The CDI will build on the results of the workshop to guide its future topics, events, and funding opportunities to support an integrated science capacity for the USGS.

Open-File Report

Oyster model inventory: Identifying critical data and modeling approaches to support restoration of oyster reefs in coastal U.S. Gulf of Mexico waters

Executive Summary Along the coast of the U.S. Gulf of Mexico, the eastern oyster ( Crassostrea virginica ) plays important ecological and economic roles. Commercial landings from this region account for more than 50 percent of all U.S. landings; these oyster reefs also provide varied ecosystem services, including nursery habitat for many fish and macroinvertebrate species, shoreline protection, and water-quality maintenance. Declining trends in both total oyster production and functional reef area across this region have spurred investment in restoration of oyster resources, with specific calls for restoration projects to develop a network of reefs and identify broodstock and sanctuary reef restoration sites. Decision making related to restoration and establishment of a network of oyster reefs in the Gulf of Mexico requires information on both the environment and the effects of the environment on the oyster life cycle (including larval movement, survival, oyster recruitment, reproduction, growth, and mortality). Here, we examined the current state of data and model development in this region with the goal of providing an overview of oyster modeling approaches and an inventory of available data and existing oyster models. This report is meant to provide an overview to managers for understanding existing efforts and identify a path forward to most efficiently inform oyster resource management and restoration planning in moving from a single reef management approach to a reef network management approach. Numerous models related to some aspect of the oyster life cycle have been built, calibrated, and validated for various Gulf of Mexico estuaries over the last few decades (over 30 models identified). These models, which could inform site restoration, can be classified into four approaches: (1) oyster Habitat Suitability Index (HSI) models; (2) larval transport models; (3) on-reef oyster models that may include oyster growth, mortality and reproduction, and substrate persistence; and (4) coupled larval transport on-reef metapopulation models that simulate the entire oyster life cycle. The data requirements, model complexity and assumptions, and transferability vary by approach. Specifically, some approaches may offer greater accessibility, flexibility, and transferability spatially or temporally, with minimal data input, but only provide broad information to support site selection. In contrast, other approaches may require significant site-specific data for their construction and validation but may provide more accurate and location-specific data to support site selection for broodstock reefs. Regardless of modeling approach used, data on environmental drivers, such as salinity, water temperature, or water flow impacting oyster metabolism and movement, are required at appropriate spatial and temporal scales. While numerous data collection platforms, environmental models, and research products exist within Gulf of Mexico estuaries to provide important environmental data to use as drivers in the oyster models, significant variability in temporal and spatial coverage of the data, and variation in the availability of future condition models, exists across estuaries. This variation influences the spatial and temporal scales at which oyster models may be developed and impacts the calibration and validation of the oyster models within a given estuary, affecting its potential ability to address specific management or restoration questions. While multiple modeling approaches exist for informing site selection of broodstock or sanctuary oyster reefs, the development, calibration, and validation of a single modeling platform presents the most efficient, transferable, and useful tool for managers across the Gulf of Mexico. The development of a single modeling platform would involve using standardized input variables, governing equations, and assumptions for the modeled oyster processes and outputs, and for standardized calibration and validation procedures that could be applied within each estuary. The differences among estuary applications would require substituting only estuary-specific environmental data, and calibrating and validating the modeling approach with local oyster data. Two modeling approaches likely to be useful include (1) development of a general geospatial HSI modeling framework that could be applied consistently across estuaries and (2) a mechanistic coupled larval transport on-reef metapopulation model requiring only estuarine specific calibration and hydrodynamic models. Both approaches benefit from existing work across multiple Gulf of Mexico estuaries and could provide valuable support for oyster restoration, but may differ in their ability to address specific questions related to oyster restoration. HSI models specifically guide restoration practitioners in determining suitable habitat based on available data. The HSI approach, while currently more widely used and accessible, requires more development of larval suitability and larval input and output components in order to inform reef connectivity. A metapopulation approach considering the full oyster life cycle that simulates both on-reef oyster growth, mortality, reproduction, substrate persistence, and larval transport (ideally with larval growth and mortality) would provide the greatest detail and level of understanding but requires significant up-front investment. The larval oyster model and on-reef oyster model are usually developed independently for systems, although the two approaches can be coupled to represent the entire oyster life cycle in order to characterize and assess a reef metapopulation. This approach may be less accessible and much more data-intensive, however, and it requires some expertise to run and apply to inform oyster resource management. Ultimately, the development of single modeling platforms for each of these approaches would provide flexible tools applicable across all Gulf of Mexico oyster supporting estuaries. By using a single platform for model development, testing, calibrating and validating, and evaluation of modeled future scenarios, oyster restoration scientists and managers would not only be able to examine different scenario outcomes within a single estuary, but could also have comparable modeled results to evaluate potential outcomes, across estuaries and regions, that are not confounded by varying modeled data inputs, governing equations, assumptions, or user judgement.

Alabama, Florida, Louisiana, Mississippi, Texas

Compilation and evaluation of data used to identify groundwater sources under the direct influence of surface water in Pennsylvania

A study was conducted to compile and evaluate data used to identify groundwater sources that are under the direct influence of surface water (GUDI) in Pennsylvania. In the early 1990s, the Pennsylvania Department of Environmental Protection (PADEP) implemented the Surface Water Identification Protocol (SWIP) for the identification of GUDI sources. Since the establishment of the SWIP, PADEP has classified more than 500 individual sources across Pennsylvania as GUDI, but Pennsylvania’s complex geology and physiography provide a challenge for a uniform method of GUDI determination. Components used in this study to compile and evaluate data associated with GUDI determination include: (1) a preliminary review of file information for 43 public water-supply wells, (2) quality control and addition of data to PADEP’s database for public water-supply systems to prepare data for analysis, and (3) exploratory evaluation of existing GUDI sources in the database with respect to hydrogeologic and source-construction characteristics that are currently utilized in the assessment methodology. Case files for 43 wells from PADEP’s Northcentral and Southcentral regions were reviewed to: (1) provide a better understanding of how the SWIP was applied in practice, (2) verify and compile missing data, and (3) find additional attributes not previously available that might explain a well’s categorization as GUDI. Review of file information showed that the SWIP outlined in PADEP technical guidance was usually followed, but for some sources, the GUDI determination was more complex and could not be easily summarized. Data compiled for study analyses provided by PADEP include source data derived from public water-supply system case files, a source-information database for public water-supply systems, and Microscopic Particulate Analysis (MPA) results and associated water-quality data for public water-supply system groundwater sources. Data from the Pennsylvania Drinking Water Information System (PADWIS) , which is PADEP’s database for public water-supply systems, were also used for this study. The PADWIS database originally included data for 12,147 groundwater sources (11,812 groundwater sources not under the direct influence of surface water (non-GUDI) wells and 335 GUDI wells). A subset (4,018 wells consisting of 3,842 non-GUDI wells and 175 GUDI wells) of the PADWIS database was created for an analysis and includes only community wells evaluated in accordance with the SWIP. MPA results for 631 community and noncommunity wells were compiled, along with associated water-quality data (alkalinity, chloride, Escherichia coli , fecal coliform, nitrate, pH, sodium, specific conductance, sulfate, total coliform, total dissolved solids, total residue, and turbidity) populated from the PADEP Bureau of Laboratories Sample Information System. Data compiled from sources other than PADEP include spatial data, both naturogenic (for example, average precipitation or distance to closest hydrologic feature) and anthropogenic (for example, percentage of developed or agricultural land cover within a specific vicinity of a public water-supply system well) data representing spatially derived variables. Comparison among wells in the PADWIS dataset subset using the nonparametric Kruskal-Wallis test showed that GUDI wells had significantly older median construction years, shallower depths, and static water levels closer to the land surface than non-GUDI wells and that carbonate aquifers had the highest percentages of wells designated as GUDI (12 percent; 57 wells). Further comparison of wells in the PADWIS database subset using the Spearman’s rho monotonic correlation test illustrated that public water-supply wells designated as GUDI largely occur in unconfined aquifers and have high average yield and shallow static water levels. Assessment of the MPA database subset using the Kruskal-Wallis test showed wells with MPA total risk-factor scores that exceeded zero had older median construction years and shallower casing depths than wells with MPA total risk-factor scores of zero and that carbonate aquifers had the highest percentages of wells with MPA total risk-factor scores exceeding zero (30 percent; 63 wells). Spearman’s rho correlations showed that wells completed in aquifers with depths to major water-bearing zones closer to the land-surface had higher total risk-factor scores resulting from MPA samples. Based on the results of the analyses described in this report, broad conclusions can be drawn regarding site-specific well characteristics as well as anthropogenic and naturogenic factors that could be responsible for a well being designated as GUDI, but the accuracy of these results is dependent on the quality of the data being analyzed. Ultimately, study results serve as an added resource for initial desktop screening of wells to determine if additional site-specific investigation is warranted and underscore the need for field evaluation.

Pennsylvania