Geology ReportsSearch

USGS · 70254127

State of the data: Assessing the FAIRness of USGS data

Abstract

In response to recent shifts towards open science that emphasize transparency, reproducibility, and access to research data, the US Geological Survey (USGS) conducted a study to assess the degree to which USGS data assets meet the FAIR data principles (Findable, Accessible, Interoperable, and Reusable). The USGS designed and applied a methodology for quantitative analysis of FAIR characteristics. A new rubric was derived from a crosswalk of existing FAIR evaluation frameworks and customized for the USGS. The rubric, consisting of 62 yes/no questions, was applied to 392 metadata records of USGS data products published between 1987 and 2022. Results were analyzed to show which FAIR characteristics were most and least present in the metadata and how these scores changed after the implementation of data policy requirements in 2016. Aggregated scores showed specific areas of strength and needed improvements. The greatest increases in FAIR scores over time were for elements that were required by new data policies, especially in the ‘Findable’ category. Based on the results, this paper presents strategies to further improve USGS alignment with FAIR. The suggested strategies are organized in four key areas: USGS data repository characteristics, training and communities of practice, data management policy considerations, and metadata standards, tools, and best practices.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vivian B. Hutchison, Tamar Norkin, Lisa Zolly, Leslie Hsu. 2024-04-26. State of the data: Assessing the FAIRness of USGS data. https://doi.org/10.5334/dsj-2024-022

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Are researchers citing their data? A case study from the U.S. Geological Survey

Data citation promotes accessibility and discoverability of data through measures carried out by researchers, publishers, repositories, and the scientific community. This paper examines how a data citation workflow has been implemented by the U.S. Geological Survey (USGS) by evaluating publication and data linkages. Two different methods were used to identify data citations: examining publication structural metadata and examining the full text of the publication. A growing number of USGS researchers are complying with publisher data sharing policies aimed to capture data citation information in a standardized way within associated publications. However, inconsistencies in how data citation information is documented in publications has limited the accessibility and discoverability of the data. This paper demonstrates how organizational evaluations of publication and data linkages can be used to identify obstacles in advancing data citation efforts and improve data citation workflows.

Data Science Journal

Developing criteria to establish Trusted Digital Repositories

This paper details the drivers, methods, and outcomes of the U.S. Geological Survey’s quest to establish criteria by which to judge its own digital preservation resources as Trusted Digital Repositories. Drivers included recent U.S. legislation focused on data and asset management conducted by federal agencies spending $100M USD or more annually on research activities. The methods entailed seeking existing evaluation criteria from national and international organizations such as International Standards Organization (ISO), U.S. Library of Congress, and Data Seal of Approval upon which to model USGS repository evaluations. Certification, complexity, cost, and usability of existing evaluation models were key considerations. The selected evaluation method was derived to allow the repository evaluation process to be transparent, understandable, and defensible; factors that are critical for judging competing, internal units. Implementing the chosen evaluation criteria involved establishing a cross-agency, multi-disciplinary team that interfaced across the organization.

Data Science Journal

Post-disaster supply chain interdependent critical infrastructure system restoration: A review of data necessary and available for modeling

The majority of restoration strategies in the wake of large-scale disasters have focused on short-term emergency response solutions. Few consider medium- to long-term restoration strategies to reconnect urban areas to national supply chain interdependent critical infrastructure systems (SCICI). These SCICI promote the effective flow of goods, services, and information vital to the economic vitality of an urban environment. To re-establish the connectivity that has been broken during a disaster between the different SCICI, relationships between these systems must be identified, formulated, and added to a common framework to form a system-level restoration plan. To accomplish this goal, a considerable collection of SCICI data is necessary. The aim of this paper is to review what data are required for model construction, the accessibility of these data, and their integration with each other. While a review of publicly available data reveals a dearth of real-time data to assist modeling long-term recovery following an extreme event, a significant amount of static data does exist and these data can be used to model the complex interdependencies needed. For the sake of illustration, a particular SCICI (transportation) is used to highlight the challenges of determining the interdependencies and creating models capable of describing the complexity of an urban environment with the data publicly available. Integration of such data as is derived from public domain sources is readily achieved in a geospatial environment, after all geospatial infrastructure data are the most abundant data source and while significant quantities of data can be acquired through public sources, a significant effort is still required to gather, develop, and integrate these data from multiple sources to build a complete model. Therefore, while continued availability of high quality, public information is essential for modeling efforts in academic as well as government communities, a more streamlined approach to a real-time acquisition and integration of these data is essential.

Data Science Journal