Geology ReportsSearch

Geology topics

Evan B. Goldstein

Publications and source records attributed to Evan B. Goldstein.

7 recordsLinked to original sources

A 1.2 billion pixel human-labeled dataset for data-driven classification of coastal environments

The world’s coastlines are spatially highly variable, coupled-human-natural systems that comprise a nested hierarchy of component landforms, ecosystems, and human interventions, each interacting over a range of space and time scales. Understanding and predicting coastline dynamics necessitates frequent observation from imaging sensors on remote sensing platforms. Machine Learning models that carry out supervised (i.e., human-guided) pixel-based classification, or image segmentation, have transformative applications in spatio-temporal mapping of dynamic environments, including transient coastal landforms, sediments, habitats, waterbodies, and water flows. However, these models require large and well-documented training and testing datasets consisting of labeled imagery. We describe “Coast Train,” a multi-labeler dataset of orthomosaic and satellite images of coastal environments and corresponding labels. These data include imagery that are diverse in space and time, and contain 1.2 billion labeled pixels, representing over 3.6 million hectares. We use a human-in-the-loop tool especially designed for rapid and reproducible Earth surface image segmentation. Our approach permits image labeling by multiple labelers, in turn enabling quantification of pixel-level agreement over individual and collections of images.

Scientific Data

A reproducible and reusable pipeline for segmentation of geoscientific imagery

Segmentation of Earth science imagery is an increasingly common task. Among modern techniques that use Deep Learning, the UNet architecture has been shown to be a reliable for segmenting a range of imagery. We developed software–Segmentation Gym–to implement a data-model pipeline for segmentation of scientific imagery using a family of UNet models. With an existing set of imagery and labels, the software uses a single configuration file that handles data set creation, as well as model setup and model training. Key benefits of this software are (a) the focus on reproducible data set creation and modeling, and (b) the ability for quick model experimentation through changes to a configuration file. Quick experimentation permits researchers to prototype different model architectures, sizes, and adjust common hyperparameters to find a suitable model. We demonstrate the use of the software using a data set of 419 labeled Landsat-8 scenes of coastal environments and compare results across two model architectures, five model sizes, and three loss functions. This demonstration highlights that our software enables rapid, reproducible experimentation to determine optimal hyperparameters for specific data sets and research questions.

Earth and Space Science

Human-in-the-Loop segmentation of earth surface imagery

Segmentation, or the classification of pixels (grid cells) in imagery, is ubiquitously applied in the natural sciences. Manual methods are often prohibitively time-consuming, especially those images consisting of small objects and/or significant spatial heterogeneity of colors or textures. Labeling complicated regions of transition that in Earth surface imagery are represented by collections of mixed-pixels, -textures, and -spectral signatures, can be especially error-prone because it is difficult to reliably unmix, identify and delineate consistently. However, the success of supervised machine learning (ML) approaches is entirely dependent on good label data. We describe a fast, semi-automated, method for interactive segmentation of N-dimensional (x, y, N) images into two-dimensional (x, y) label images. It uses human-in-the-loop ML to achieve consensus between the labeler and a model in an iterative workflow. The technique is reproducible; the sequence of decisions made by human labeler and ML algorithms can be encoded to file, so the entire process can be played back and new outputs generated with alternative decisions and/or algorithms. We illustrate the scientific potential of segmentation of imagery of diverse settings and image types using six case studies from river, estuarine, and open coast environments. These photographic and non-photographic imagery consist of 1- and 3-bands on regular and irregular grids ranging from centimeters to tens of meters. We demonstrate high levels of agreement in label images generated by several labelers on the same imagery, and make suggestions to achieve consensus and measure uncertainty, ideal for widespread application in training supervised ML for image segmentation.

Earth and Space Science

Labeling poststorm coastal imagery for machine learning: Measurement of interrater agreement

Classifying images using supervised machine learning (ML) relies on labeled training data—classes or text descriptions, for example, associated with each image. Data-driven models are only as good as the data used for training, and this points to the importance of high-quality labeled data for developing a ML model that has predictive skill. Labeling data is typically a time-consuming, manual process. Here, we investigate the process of labeling data, with a specific focus on coastal aerial imagery captured in the wake of hurricanes that affected the Atlantic and Gulf Coasts of the United States. The imagery data set is a rich observational record of storm impacts and coastal change, but the imagery requires labeling to render that information accessible. We created an online interface that served labelers a stream of images and a fixed set of questions. A total of 1,600 images were labeled by at least two or as many as seven coastal scientists. We used the resulting data set to investigate interrater agreement: the extent to which labelers labeled each image similarly. Interrater agreement scores, assessed with percent agreement and Krippendorff's alpha, are higher when the questions posed to labelers are relatively simple, when the labelers are provided with a user manual, and when images are smaller. Experiments in interrater agreement point toward the benefit of multiple labelers for understanding the uncertainty in labeling data for machine learning research.

Earth and Space Science

Blind testing of shoreline evolution models

Beaches around the world continuously adjust to daily and seasonal changes in wave and tide conditions, which are themselves changing over longer time-scales. Different approaches to predict multi-year shoreline evolution have been implemented; however, robust and reliable predictions of shoreline evolution are still problematic even in short-term scenarios (shorter than decadal). Here we show results of a modelling competition, where 19 numerical models (a mix of established shoreline models and machine learning techniques) were tested using data collected for Tairua beach, New Zealand with 18 years of daily averaged alongshore shoreline position and beach rotation (orientation) data obtained from a camera system. In general, traditional shoreline models and machine learning techniques were able to reproduce shoreline changes during the calibration period (1999–2014) for normal conditions but some of the model struggled to predict extreme and fast oscillations. During the forecast period (unseen data, 2014–2017), both approaches showed a decrease in models’ capability to predict the shoreline position. This was more evident for some of the machine learning algorithms. A model ensemble performed better than individual models and enables assessment of uncertainties in model architecture. Research-coordinated approaches (e.g., modelling competitions) can fuel advances in predictive capabilities and provide a forum for the discussion about the advantages/disadvantages of available models.

Scientific Reports

Building back bigger in hurricane strike zones

Despite decades of regulatory efforts in the United States to decrease vulnerability in developed coastal zones, exposure of residential assets to hurricane damage is increasing — even in places where hurricanes have struck before. Comparing plan-view footprints of individual residential buildings before and long after major hurricane strikes, we find a systematic pattern of ‘building back bigger’ among renovated and new properties.

Nature Sustainability

Indications of a positive feedback between coastal development and beach nourishment

Beach nourishment, a method for mitigating coastal storm damage or chronic erosion by deliberately replacing sand on an eroded beach, has been the leading form of coastal protection in the U.S. for four decades. However, investment in hazard protection can have the unintended consequence of encouraging development in places especially vulnerable to damage. In a comprehensive, parcel-scale analysis of all shorefront single-family homes in the state of Florida, we find that houses in nourishing zones are significantly larger and more numerous than in non-nourishing zones. The predominance of larger homes in nourishing zones suggests a positive feedback between nourishment and development that is compounding coastal risk in zones already characterized by high vulnerability.

Earth's Future