Geology Reports⌕ Search

SEARCH · Geology Reports

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

An integrated modeling approach to estimating Gunnison Sage-Grouse population dynamics: Combining index and demographic data

Evaluation of population dynamics for rare and declining species is often limited to data that are sparse and/or of poor quality. Frequently, the best data available for rare bird species are based on large‐scale, population count data. These data are commonly based on sampling methods that lack consistent sampling effort, do not account for detectability, and are complicated by observer bias. For some species, short‐term studies of demographic rates have been conducted as well, but the data from such studies are typically analyzed separately. To utilize the strengths and minimize the weaknesses of these two data types, we developed a novel Bayesian integrated model that links population count data and population demographic data through population growth rate ( λ ) for Gunnison sage‐grouse ( Centrocercus minimus ). The long‐term population index data available for Gunnison sage‐grouse are annual (years 1953–2012) male lek counts. An intensive demographic study was also conducted from years 2005 to 2010. We were able to reduce the variability in expected population growth rates across time, while correcting for potential small sample size bias in the demographic data. We found the population of Gunnison sage‐grouse to be variable and slightly declining over the past 16 years.

Colorado, Utah↗

Classifying behavior from short-interval biologging data: An example with GPS tracking of birds

Recent advances in digital data collection have spurred accumulation of immense quantities of data that have potential to lead to remarkable ecological insight, but that also present analytic challenges. In the case of biologging data from birds, common analytical approaches to classifying movement behaviors are largely inappropriate for these massive data sets. We apply a framework for using K -means clustering to classify bird behavior using points from short time interval GPS tracks. K -means clustering is a well-known and computationally efficient statistical tool that has been used in animal movement studies primarily for clustering segments of consecutive points. To illustrate the utility of our approach, we apply K -means clustering to six focal variables derived from GPS data collected at 1–11 s intervals from free-flying bald eagles ( Haliaeetus leucocephalus ) throughout the state of Iowa, USA. We illustrate how these data can be used to identify behaviors and life-stage- and age-related variation in behavior. After filtering for data quality, the K -means algorithm identified four clusters in >2 million GPS telemetry data points. These four clusters corresponded to three movement states: ascending, flapping, and gliding flight; and one non-moving state: perching. Mapping these states illustrated how they corresponded tightly to expectations derived from natural history observations; for example, long periods of ascending flight were often followed by long gliding descents, birds alternated between flapping and gliding flight. The K -means clustering approach we applied is both an efficient and effective mechanism to classify and interpret short-interval biologging data to understand movement behaviors. Furthermore, because it can apply to an abundance of very short, irregular, and high-dimensional movement data, it provides insight into small-scale variation in behavior that would not be possible with many other analytical approaches.

Ecology and Evolution↗

Errors in aerial survey count data: Identifying pitfalls and solutions

Accurate estimates of animal abundance are essential for guiding effective management, and poor survey data can produce misleading inferences. Aerial surveys are an efficient survey platform, capable of collecting wildlife data across large spatial extents in short timeframes. However, these surveys can yield unreliable data if not carefully executed. Despite a long history of aerial survey use in ecological research, problems common to aerial surveys have not yet been adequately resolved. Through an extensive review of the aerial survey literature over the last 50 years, we evaluated how common problems encountered in the data (including nondetection, counting error, and species misidentification) can manifest, the potential difficulties conferred, and the history of how these challenges have been addressed. Additionally, we used a double-observer case study focused on waterbird data collected via aerial surveys and an online group (flock) counting quiz to explore the potential extent of each challenge and possible resolutions. We found that nearly three quarters of the aerial survey methodology literature focused on accounting for nondetection errors, while issues of counting error and misidentification were less commonly addressed. Through our case study, we demonstrated how these challenges can prove problematic by detailing the extent and magnitude of potential errors. Using our online quiz, we showed that aerial observers typically undercount group size and that the magnitude of counting errors increases with group size. Our results illustrate how each issue can act to bias inferences, highlighting the importance of considering individual methods for mitigating potential problems separately during survey design and analysis. We synthesized the information gained from our analyses to evaluate strategies for overcoming the challenges of using aerial survey data to estimate wildlife abundance, such as digital data collection methods, pooling species records by family, and ordinal modeling using binned data. Recognizing conditions that can lead to data collection errors and having reasonable solutions for addressing errors can allow researchers to allocate resources effectively to mitigate the most significant challenges for obtaining reliable aerial survey data.

Alabama, Florida, Louisiana, Mississippi, Texas↗

Missing data in ecology: Syntheses, clarifications, and considerations

In ecology and related sciences, missing data are common and occur in a variety of different contexts. When missing data are not handled properly, subsequent statistical estimates tend to be biased, inefficient, and lack proper confidence interval coverage. Missing data are often grouped into three categories: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). We review each category and compare their benefits and drawbacks. We review several approaches to handling missing data including complete case analysis, imputation, inverse probability weighting, and data augmentation. We clarify what types of variables should accompany imputation methods and how those variables are influenced by the analysis methods. Additionally, we discuss missing data that lack a formal basis for measurement and hence are fundamentally different from MCAR, MAR, and MNAR missing data. Throughout, we introduce concepts and numeric examples using both simulated data and data from the United States Environmental Protection Agency's 2016 National Wetland Condition Assessment. We conclude by providing five considerations for ecologists and other scientists handling missing data.

Ecological Monographs↗

Incorporating citizen science data in spatially explicit integrated population models

Information about population abundance, distribution, and demographic rates is critical for understanding a species’ ecology and for effective conservation and management. To collect data over large spatial and temporal extents for such inferences, especially for species with low densities or wide distributions, citizen science can be an efficient approach. Integrated models have also emerged as an important methodology to estimate population parameters by combining multiple types of data, including citizen science data. We developed a spatially explicit integrated model that combines opportunistically collected presence–absence (PA) data, commonly collected in citizen science efforts, with systematically collected spatial capture–recapture (SCR) data, which are often limited to small spatial and temporal extents. We conducted single and multi‐season simulations with parameters informed by North American black bear ( Ursus americanus ) populations, to evaluate the influence of varying amounts of opportunistic PA data collected at larger spatial and temporal extents on the estimation of population‐level parameters. Integrating opportunistic PA data increased the precision and accuracy of posterior estimates of abundance, and survival and recruitment rates. In some cases, adding PA locations improved abundance estimates more than increasing PA detection probability. Posterior estimates were as precise and unbiased as when higher quality, but sparse, SCR data were available. We also applied the integrated model to SCR and citizen science PA data collected on black bears in New York, with results consistent with our simulations. Our findings indicate that citizen science in integrated models can be a cost‐efficient way to improve estimates of population parameters and increase the spatiotemporal extent of inference. Continued developments with integrated models and citizen science data will offer additional ways to improve our understanding of population structure and demographics.

Ecology↗

A metadata checklist and data formatting guidelines to make eDNA FAIR (Findable, Accessible, Interoperable and Reusable)

The success of environmental DNA (eDNA) approaches for species detection has revolutionized biodiversity monitoring and distribution mapping. Targeted eDNA amplification approaches, such as quantitative PCR, have improved our understanding of species distribution, and metabarcoding-based approaches have enabled biodiversity assessment at unprecedented scales and taxonomic resolution. eDNA datasets, however, are often scattered across repositories with inconsistent formats, varying access restrictions, and inadequate metadata; this limits their interoperation, reuse, and overall impact. Adopting FAIR (Findable, Accessible, Interoperable, and Reusable) data practices with eDNA data can transform the monitoring of biodiversity and individual species and support data-driven biodiversity management across broad scales. FAIR practices remain underdeveloped in the eDNA community, partly due to gaps in adapting existing vocabularies, such as Darwin Core (DwC) and Minimum Information about any (x) Sequence (MIxS), to eDNA-specific needs and workflows. To address these challenges, we propose a comprehensive FAIR eDNA (FAIRe) Metadata Checklist, which integrates existing data standards and introduces new terms tailored to eDNA workflows. Metadata are systematically linked to both raw data (e.g., metabarcoding sequences, Ct/Cq values of targeted qPCR assays) and derived biological observations (e.g., Amplicon Sequence Variant (ASV)/Operational Taxonomic Unit (OTU) tables, species presence/absence). Along with formatting guidelines, tools, templates, and example datasets, we introduce a standardized, ready-to-use approach for FAIR eDNA practices. Through broad collaboration, we seek to integrate these guidelines into established biodiversity and molecular data standards, promote journal data policies, and foster user-driven improvements and uptake of FAIR practices among eDNA data producers. In proposing this standardized approach and developing a long-term plan with key databases and data standard organizations, the goal is to enhance accessibility, maximize reuse, and elevate the scientific impact of these valuable biodiversity data resources.

Environmental DNA↗

Data worth and prediction uncertainty for pesticide transport and fate models in Nebraska and Maryland, United States

BACKGROUND Complex environmental models are frequently extrapolated to overcome data limitations in space and time, but quantifying data worth to such models is rarely attempted. The authors determined which field observations most informed the parameters of agricultural system models applied to field sites in Nebraska (NE) and Maryland (MD), and identified parameters and observations that most influenced prediction uncertainty. RESULTS The standard error of regression of the calibrated models was about the same at both NE (0.59) and MD (0.58), and overall reductions in prediction uncertainties of metolachlor and metolachlor ethane sulfonic acid concentrations were 98.0 and 98.6% respectively. Observation data groups reduced the prediction uncertainty by 55–90% at NE and by 28–96% at MD. Soil hydraulic parameters were well informed by the observed data at both sites, but pesticide and macropore properties had comparatively larger contributions after model calibration. CONCLUSIONS Although the observed data were sparse, they substantially reduced prediction uncertainty in unsampled regions of pesticide breakthrough curves. Nitrate evidently functioned as a surrogate for soil hydraulic data in well-drained loam soils conducive to conservative transport of nitrogen. Pesticide properties and macropore parameters could most benefit from improved characterization further to reduce model misfit and prediction uncertainty. RESULTS: The standard error of regression of the calibrated models was about the same at both NE (0.59) and MD (0.58), and overall reductions in prediction uncertainties of metolachlor and metolachlor ethane sulfonic acid concentrations were 98.0 and 98.6% respectively. Observation data groups reduced the prediction uncertainty by 55–90% at NE and by 28–96% at MD. Soil hydraulic parameters were well informed by the observed data at both sites, but pesticide and macropore properties had comparatively larger contributions after model calibration. CONCLUSIONS: Although the observed data were sparse, they substantially reduced prediction uncertainty in unsampled regions of pesticide breakthrough curves. Nitrate evidently functioned as a surrogate for soil hydraulic data in well-drained loam soils conducive to conservative transport of nitrogen. Pesticide properties and macropore parameters could most benefit from improved characterization further to reduce model misfit and prediction uncertainty.

Maryl;Nebraska↗

Integrated groundwater data management

The goal of a data manager is to ensure that data is safely stored, adequately described, discoverable and easily accessible. However, to keep pace with the evolution of groundwater studies in the last decade, the associated data and data management requirements have changed significantly. In particular, there is a growing recognition that management questions cannot be adequately answered by single discipline studies. This has led a push towards the paradigm of integrated modeling, where diverse parts of the hydrological cycle and its human connections are included. This chapter describes groundwater data management practices, and reviews the current state of the art with enterprise groundwater database management systems. It also includes discussion on commonly used data management models, detailing typical data management lifecycles. We discuss the growing use of web services and open standards such as GWML and WaterML2.0 to exchange groundwater information and knowledge, and the need for national data networks. We also discuss cross-jurisdictional interoperability issues, based on our experience sharing groundwater data across the US/Canadian border. Lastly, we present some future trends relating to groundwater data management.

Book chapter↗

Comparison of temporal trends in ambient and compliance trace element and PCB data in pool 2 of the Mississippi River, USA, 1985-1995

The Intergovernmental Task Force on Monitoring has suggested studies on ambient (in-stream) and compliance (wastewater) data to determine if monitoring can be reduced locally or nationally. The similarity in temporal trends between retrospective ambient and compliance water-quality data collected from Pool 2 of the Mississippi River, USA, was determined for 1985–1995. Constituents studied included the following trace elements: arsenic (As), cadmium (Cd), chromium (Cr), hexavalent chromium (Cr61), copper (Cu), lead (Pb), mercury (Hg), nickel (Ni), selenium (Se), zinc (Zn), and polychlorinated biphenyls (PCBs). Water-column, bed-sediment, and fish-tissue (fillets) data collected by five government agencies comprised the ambient data set; effluent data from five registered facilities comprised the compliance data set. The nonparametric MannKendall trend test indicated that 33% of temporal trends in all data were statistically significant (P , 0.05). Possible reasons for this were low sample sizes, and a high percentage of samples below the analytical detection limit. Trends in compliance data were more distinct; most trace elements decreased significantly, probably due to improvements in wastewater treatment. Seven trace elements (Cr, Cd, Cu, Pb, Hg, Ni, and Zn) had statistically significant decreases in wastewater and portions of either or both ambient water and bed sediment. No trends were found in fish tissue. Inconsistency in trends between ambient and compliance data were often found for individual constituents, making overall similarity between the data sets difficult to determine. Logistical differences in monitoring programs, such as varying field and laboratory methods among agencies, made it difficult to assess ambient temporal trends.

Minnesota↗

Recent data and information system initiatives for remotely sensed measurements of the land surface

As part of the International Satellite Land Satellite Climatology Program (ISLSCP) Workshop on Remote Sensing of the Land Surface for Studies of Global Change, five invited presentations were given on recent data initiatives relevant to the ISLSCP community. The presentations are summarized in this paper along with some observations by the authors on data systems for the land sciences community. The invited presentations are by no means all inclusive but were selected as examples of current data activities, representing a range of topics associated with data for land sciences including: the generation of global and local scale data sets, the reworking of historical data sets, new data initiatives and some programmatic aspects of land data base development. This paper serves to provide information on these data initiatives and to air some of the issues concerning land science data systems that were raised at the meeting.

Remote Sensing of Environment↗

Small values in big data: The continuing need for appropriate metadata

Compiling data from disparate sources to address pressing ecological issues is increasingly common. Many ecological datasets contain left-censored data – observations below an analytical detection limit. Studies from single and typically small datasets show that common approaches for handling censored data — e.g., deletion or substituting fixed values — result in systematic biases. However, no studies have explored the degree to which the documentation and presence of censored data influence outcomes from large, multi-sourced datasets. We describe left-censored data in a lake water quality database assembled from 74 sources and illustrate the challenges of dealing with small values in big data, including detection limits that are absent, range widely, and show trends over time. We show that substitutions of censored data can also bias analyses using ‘big data’ datasets, that censored data can be effectively handled with modern quantitative approaches, but that such approaches rely on accurate metadata that describe treatment of censored data from each source.

Ecological Informatics↗

A working environment for digital planetary data processing and mapping using ISIS and GRASS GIS

Since the beginning of planetary exploration, mapping has been fundamental to summarize observations returned by scientific missions. Sensor-based mapping has been used to highlight specific features from the planetary surfaces by means of processing. Interpretative mapping makes use of instrumental observations to produce thematic maps that summarize observations of actual data into a specific theme. Geologic maps, for example, are thematic interpretative maps that focus on the representation of materials and processes and their relative timing. The advancements in technology of the last 30 years have allowed us to develop specialized systems where the mapping process can be made entirely in the digital domain. The spread of networked computers on a global scale allowed the rapid propagation of software and digital data such that every researcher can now access digital mapping facilities on his desktop. The efforts to maintain planetary missions data accessible to the scientific community have led to the creation of standardized digital archives that facilitate the access to different datasets by software capable of processing these data from the raw level to the map projected one. Geographic Information Systems (GIS) have been developed to optimize the storage, the analysis, and the retrieval of spatially referenced Earth based environmental geodata; since the last decade these computer programs have become popular among the planetary science community, and recent mission data start to be distributed in formats compatible with these systems. Among all the systems developed for the analysis of planetary and spatially referenced data, we have created a working environment combining two software suites that have similar characteristics in their modular design, their development history, their policy of distribution and their support system. The first, the Integrated Software for Imagers and Spectrometers (ISIS) developed by the United States Geological Survey, represents the state of the art for processing planetary remote sensing data, from the raw unprocessed state to the map projected product. The second, the Geographic Resources Analysis Support System (GRASS) is a Geographic Information System developed by an international team of developers, and one of the core projects promoted by the Open Source Geospatial Foundation (OSGeo). We have worked on enabling the combined use of these software systems throughout the set-up of a common user interface, the unification of the cartographic reference system nomenclature and the minimization of data conversion. Both software packages are distributed with free open source licenses, as well as the source code, scripts and configuration files hereafter presented. In this paper we describe our work done to merge these working environments into a common one, where the user benefits from functionalities of both systems without the need to switch or transfer data from one software suite to the other one. Thereafter we provide an example of its usage in the handling of planetary data and the crafting of a digital geologic map. ?? 2010 Elsevier Ltd. All rights reserved.

Conference Paper↗

Developing a national stream morphology data exchange: needs, challenges, and opportunities

Stream morphology data, primarily consisting of channel and foodplain geometry and bed material size measurements, historically have had a wide range of applications and uses including culvert/ bridge design, rainfall- runoff modeling, food inundation mapping (e.g., U.S. Federal Emergency Management Agency food insurance studies), climate change studies, channel stability/sediment source investigations, navigation studies, habitat assessments, and landscape change research. The need for stream morphology data in the United States, and thus the quantity of data collected, has grown substantially over the past 2 decades because of the expanded interests of resource management agencies in watershed management and restoration. The quantity of stream morphology data collected has also increased because of state-of-the-art technologies capable of rapidly collecting high-resolution data over large areas with heretofore unprecedented precision. Despite increasing needs for and the expanding quantity of stream morphology data, neither common reporting standards nor a central data archive exist for storing and serving these often large and spatially complex data sets. We are proposing an open- access data exchange for archiving and disseminating stream morphology data.

Eos, Transactions, American Geophysical Union↗

Cloud-native repositories for big scientific data

Scientific data have traditionally been distributed via downloads from data server to local computer. This way of working suffers from limitations as scientific datasets grow toward the petabyte scale. A “cloud-native data repository,” as defined in this article, offers several advantages over traditional data repositories—performance, reliability, cost-effectiveness, collaboration, reproducibility, creativity, downstream impacts, and access and inclusion. These objectives motivate a set of best practices for cloud-native data repositories: analysis-ready data, cloud-optimized (ARCO) formats, and loose coupling with data-proximate computing. The Pangeo Project has developed a prototype implementation of these principles by using open-source scientific Python tools. By providing an ARCO data catalog together with on-demand, scalable distributed computing, Pangeo enables users to process big data at rates exceeding 10 GB/s. Several challenges must be resolved in order to realize cloud computing’s full potential for scientific research, such as organizing funding, training users, and enforcing data privacy requirements.

Computing in Science and Engineering↗

Atmospheric correction at AERONET locations: A new science and validation data set

This paper describes an Aerosol Robotic Network (AERONET)-based Surface Reflectance Validation Network (ASRVN) and its data set of spectral surface bidirectional reflectance and albedo based on Moderate Resolution Imaging Spectroradiometer (MODIS) TERRA and AQUA data. The ASRVN is an operational data collection and processing system. It receives 50 ?? 50 km 2 ; subsets of MODIS level 1B (L1B) data from MODIS adaptive processing system and AERONET aerosol and water-vapor information. Then, it performs an atmospheric correction (AC) for about 100 AERONET sites based on accurate radiative-transfer theory with complex quality control of the input data. The ASRVN processing software consists of an L1B data gridding algorithm, a new cloud-mask (CM) algorithm based on a time-series analysis, and an AC algorithm using ancillary AERONET aerosol and water-vapor data. The AC is achieved by fitting the MODIS top-of-atmosphere measurements, accumulated for a 16-day interval, with theoretical reflectance parameterized in terms of the coefficients of the Li SparseRoss Thick (LSRT) model of the bidirectional reflectance factor (BRF). The ASRVN takes several steps to ensure high quality of results: 1) the filtering of opaque clouds by a CM algorithm; 2) the development of an aerosol filter to filter residual semitransparent and subpixel clouds, as well as cases with high inhomogeneity of aerosols in the processing area; 3) imposing the requirement of the consistency of the new solution with previously retrieved BRF and albedo; 4) rapid adjustment of the 16-day retrieval to the surface changes using the last day of measurements; and 5) development of a seasonal backup spectral BRF database to increase data coverage. The ASRVN provides a gapless or near-gapless coverage for the processing area. The gaps, caused by clouds, are filled most naturally with the latest solution for a given pixel. The ASRVN products include three parameters of the LSRT model (kL, kG, and kV), surface albedo, normalized BRF (computed for a standard viewing geometry, VZA = 0, SZA = 45??), and instantaneous BRF (or one-angle BRF value derived from the last day of MODIS measurement for specific viewing geometry) for the MODIS 500-m bands 17. The results are produced daily at a resolution of 1 km in gridded format. We also provide a cloud mask, a quality flag, and a browse bitmap image. The ASRVN data set, including 6 years of MODIS TERRA and 1.5 years of MODIS AQUA data, is available now as a standard MODIS product (MODASRVN) which can be accessed through the Level 1 and Atmosphere Archive and Distribution System website ( http://ladsweb.nascom.nasa.gov/data/search.html). It can be used for a wide range of applications including validation analysis and science research. ?? 2006 IEEE.

IEEE Transactions on Geoscience and Remote Sensing↗

A procedure for radiometric recalibration of Landsat 5 TM reflective-band data

From the Landsat program's inception in 1972 to the present, the Earth science user community has been benefiting from a historical record of remotely sensed data. The multispectral data from the Landsat 5 (L5) Thematic Mapper (TM) sensor provide the backbone for this extensive archive. Historically, the radiometric calibration procedure for the L5 TM imagery used the detectors' response to the internal calibrator (IC) on a scene-by-scene basis to determine the gain and offset for each detector. The IC system degraded with time, causing radiometric calibration errors up to 20%. In May 2003, the L5 TM data processed and distributed by the U.S. Geological Survey (USGS) Earth Resources Observation and Science Center through the National Landsat Archive Production System (NLAPS) were updated to use a lifetime lookup-table (LUT) gain model to radiometrically calibrate TM data instead of using scene-specific IC gains. Further modification of the gain model was performed in 2007. The L5 TM data processed using IC prior to the calibration update do not benefit from the recent calibration revisions. A procedure has been developed to give users the ability to recalibrate their existing level-1 products. The best recalibration results are obtained if the work-order report that was included in the original standard data product delivery is available. However, if users do not have the original work-order report, the IC trends can be used for recalibration. The IC trends were generated using the radiometric gain trends recorded in the NLAPS database. This paper provides the details of the recalibration procedure for the following: 1) data processed using IC where users have the work-order file; 2) data processed using IC where users do not have the work-order file; 3) data processed using prelaunch calibration parameters; and 4) data processed using the previous version of the LUT (e.g., LUT03) that was released before April 2, 2007.

IEEE Transactions on Geoscience and Remote Sensing↗

Safari Science: Assessing the reliability of citizen science data for wildlife surveys

Protected areas are the cornerstone of global conservation, yet financial support for basic monitoring infrastructure is lacking in 60% of them. Citizen science holds potential to address these shortcomings in wildlife monitoring, particularly for resource-limited conservation initiatives in developing countries – if we can account for the reliability of data produced by volunteer citizen scientists (VCS). This study tests the reliability of VCS data vs. data produced by trained ecologists, presenting a hierarchical framework for integrating diverse datasets to assess extra variability from VCS data. Our results show that while VCS data are likely to be overdispersed for our system, the overdispersion varies widely by species. We contend that citizen science methods, within the context of East African drylands, may be more appropriate for species with large body sizes, which are relatively rare, or those that form small herds. VCS perceptions of the charisma of a species may also influence their enthusiasm for recording it. Tailored programme design (such as incentives for VCS) may mitigate the biases in citizen science data and improve overall participation. However, the cost of designing and implementing high-quality citizen science programmes may be prohibitive for the small protected areas that would most benefit from these approaches. Synthesis and applications . As citizen science methods continue to gain momentum, it is critical that managers remain cautious in their implementation of these programmes while working to ensure methods match data purpose. Context-specific tests of citizen science data quality can improve programme implementation, and separate data models should be used when volunteer citizen scientists' variability differs from trained ecologists' data. Partnerships across protected areas and between protected areas and other conservation institutions could help to cover the costs of citizen science programme design and implementation.

Journal of Applied Ecology↗

The Coastal Carbon Library and Atlas: Open source soil data and tools supporting blue carbon research and policy

Quantifying carbon fluxes into and out of coastal soils is critical to meeting greenhouse gas reduction and coastal resiliency goals. Numerous ‘blue carbon’ studies have generated, or benefitted from, synthetic datasets. However, the community those efforts inspired does not have a centralized, standardized database of disaggregated data used to estimate carbon stocks and fluxes. In this paper, we describe a data structure designed to standardize data reporting, maximize reuse, and maintain a chain of credit from synthesis to original source. We introduce version 1.0.0. of the Coastal Carbon Library, a global database of 6723 soil profiles representing blue carbon-storing systems including marshes, mangroves, tidal freshwater forests, and seagrasses. We also present the Coastal Carbon Atlas, an R-shiny application that can be used to visualize, query, and download portions of the Coastal Carbon Library. The majority (4815) of entries in the database can be used for carbon stock assessments without the need for interpolating missing soil variables, 533 are available for estimating carbon burial rate, and 326 are useful for fitting dynamic soil formation models. Organic matter density significantly varied by habitat with tidal freshwater forests having the highest density, and seagrasses having the lowest. Future work could involve expansion of the synthesis to include more deep stock assessments, increasing the representation of data outside of the U.S., and increasing the amount of data available for mangroves and seagrasses, especially carbon burial rate data. We present proposed best practices for blue carbon data including an emphasis on disaggregation, data publication, dataset documentation, and use of standardized vocabulary and templates whenever appropriate. To conclude, the Coastal Carbon Library and Atlas serve as a general example of a grassroots F.A.I.R. (Findable, Accessible, Interoperable, and Reusable) data effort demonstrating how data producers can coordinate to develop tools relevant to policy and decision-making.

Global Change Biology↗