Geology ReportsSearch

USGS · 70256583

Evaluating a tandem human-machine approach to labelling of wildlife in remote camera monitoring

Abstract

Remote cameras (“trail cameras”) are a popular tool for non-invasive, continuous wildlife monitoring, and as they become more prevalent in wildlife research, machine learning (ML) is increasingly used to automate or accelerate the labor-intensive process of labelling (i.e., tagging) photos. Human-machine hybrid tagging approaches have been shown to greatly increase tagging efficiency (i.e., time to tag a single image). However, those potential increases hinge on the extent to which an ML model makes correct vs. incorrect predictions. We performed an experiment using a ML model that produces bounding boxes around animals, people, and vehicles in remote camera imagery (MegaDetector) to consider the impact of a ML model’s performance on its ability to accelerate human labeling. Six participants tagged trail camera images collected from 12 sites in Vermont and Maine, USA (January–September 2022) using three tagging methods (one with ML bounding box assistance and two without assistance). We used a generalized linear mixed model to examine the influence of ML model performance and tagging method on tagging efficiency. We found that ML bounding boxes offer significant improvement in tagging efficiency when labelling data compared to unassisted tagging. Additionally, the time taken to label with bounding boxes was not statistically different from an unassisted tagging approach. However, we found that gains in efficiency are contingent on the ML algorithm’s performance and that incorrect ML predictions, particularly the 4.2% false positive and 3.6% false negative predictions, can slow the tagging process compared to a non-hybrid approach. These findings indicate that although practitioners usually forgo the production of bounding boxes when selecting a data labelling process due to the increased effort, ML bounding box-assisted tagging can offer an efficient method for labeling. More broadly, ML-assisted data labelling offers an opportunity to accelerate the analysis of trail camera imagery, but an assessment of the ML model’s performance can illuminate whether the hybrid-tagging approach is ultimately a help or hinderance.

Explore related subjects

90° N90° S · 180° W ← longitude → 180° E
Source-reported bounding extent: 42.729094° to 47.459686° latitude; -73.43774° to -66.950569° longitude. This indicates report coverage, not an exact sampling location. View area on OpenStreetMap.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Laurence A. Clarfeld, Alexej P.K. Sirén, Brendan M. Mulhall, Tammy L. Wilson, Elena Bernier, John Farrell, Gus Lunde, Nicole Hardy, Katherine D. Gieder, Robert Abrams, Sue Staats, Scott McLellan, Therese M. Donovan. 2023. Evaluating a tandem human-machine approach to labelling of wildlife in remote camera monitoring. https://doi.org/10.1016/j.ecoinf.2023.102257

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

A Bayesian hierarchical modeling approach for species diversity in ecology

Species diversity is the foundation of many ecological disciplines. This metric is often approximated using species richness and evenness, even though actual richness likely exceeds observations due to imperfect sampling methods. Estimating the “true” species richness, which includes identifying the number of missing species, has intrigued ecologists for decades. We adopted a parametric model that appeared in Fisher et al. (1943), which models the numbers of individuals from different species as random samples from a negative binomial distribution, and developed a Bayesian computational approach to directly estimate the distribution model parameters. The model parameters represent species abundance and evenness, and can be used to derive species richness. We evaluated our parametric approach using (1) a simulation study and (2) three historical data sets. Furthermore, we illustrated the hierarchical modeling approach to combine data from multiple parallel studies using a biannual fishery survey data set. Our parametric model formulation is computationally efficient, and the hierarchical structure facilitates embedding diversity estimation into broader application, such as assessing spatial and temporal trends in species diversity associated with environmental stressors. Additionally, because the two parameters of the negative binomial distribution model represent species abundance and evenness of a community, this parametric approach facilitates a deeper understanding of the ecological systems under study. The negative binomial distribution model works with a wide range of species frequency distribution types. As a result, our emphasis on a parametric model can help us characterize the structure of an ecosystem and provide a greater depth of ecologically meaningful information.

Ecological Informatics

Hierarchical mixture models and high-resolution monitoring data can inform siting and operational strategies to mitigate bat fatalities at wind turbines

Bats provide critical ecosystem services, but bat fatalities due to wind energy development may imperil some bat populations. Statistical models are used to estimate the total fatalities that occur based on carcasses observed during monitoring surveys. Current models often estimate fatalities aggregated across species, time, and/or turbines, but fall short of reliably informing siting and operational collision mitigation strategies that account for species-specific fatality patterns on a fine spatiotemporal scale. We developed a hierarchical mixture model for estimating species-specific covariate effects and total fatalities per species at each turbine on weekly intervals. We applied the model to a high-resolution dataset of bat carcasses found during turbine searches across nineteen wind facilities in Iowa over two years. Our model explains species-specific variation in bat fatalities at individual wind turbines according to turbine proximity to bat habitat, turbine design specifications, seasonal trends, and weather conditions such as nightly air temperature, air pressure, and wind speed. Turbines located on the edge of wind facilities had higher fatalities, and proximity to roosting and foraging habitat accounted for variation in species-specific fatality estimates. These insights into turbine placement effects can inform siting strategies. We also discovered species-specific relationships with average nightly wind speed and air temperature, among other weather conditions, that could inform operational mitigation strategies such as smart curtailment. Our model can transform observations of carcasses found during turbine searches across multiple facilities, years, and variable search efforts into estimates of total fatalities per species associated with species-specific spatial, temporal, and environmental covariate effects.

Ecological Informatics

Two-stage approach to automatic detection with machine learning for improved surveillance of the invasive Cuban treefrog

The Cuban treefrog ( Osteopilus septentrionalis ), as an invasive species in the southern United States, presents a need for effective surveillance. Automated detection expedites processing of audio data for large-scale surveillance and monitoring programs. However, current available methods commonly used for anuran species have not been sufficient to detect Cuban treefrogs. Here, we present results from a two-stage method for automated detection that employs both cross-correlation template matching and secondary supervised learning classifiers. In the first stage, audio data are screened for initial detections using template matching, in which the detections contain both true and false positives. In the second stage, the false positives are screened out using classifier algorithms. We used this method to process 139,985 audio recordings, consisting of 596,046 total minutes, collected at 13 locations in Louisiana and Florida from 2014 to 2022. From the stage 1 template matching, we detected 83,191 Cuban treefrog signals across recordings. The stage 2 machine learning model was able to identify stage 1 false positive detections with a testing accuracy of 98.46% and a testing false positive rate of 1.116%. After pruning false positive detections, a total of 20,271 individual Cuban treefrog detections remained, distributed mainly across 3 sites in an area with known presence. Locations with presumed absence had an easily verifiable number of false positive detections ( n = 109 across all other sites). The two-stage methodology utilizing both template matching and machine learning algorithms can be integrated into wildlife surveillance or monitoring programs for species with distinctive, conserved calls as an effective way to achieve sensitive species detection with a low incidence of false positives.

Florida, Louisiana