Geology ReportsSearch

USGS · 70227104

Attack of the PCR clones: Rates of clonality have little effect on RAD-seq genotype calls

Abstract

Interpretation of high-throughput sequence data requires an understanding of how decisions made during bioinformatic data processing can influence results. One source of bias that is often cited is PCR clones (or PCR duplicates). PCR clones are common in restriction site-associated sequencing (RAD-seq) data sets, which are increasingly being used for molecular ecology. To determine the influence PCR clones and the bioinformatic handling of clones have on genotyping, we evaluate four RAD-seq data sets. Data sets were compared before and after clones were removed to estimate the number of clones present in RAD-seq data, quantify how often the presence of clones in a data set causes genotype calls to change compared to when clones were removed, investigate the mechanisms that lead to genotype call changes and test whether clones bias heterozygosity estimates. Our RAD-seq data sets contained 30%–60% PCR clones, but 95% of RAD-tags had five or fewer clones. Relatively few genotypes changed once clones were removed (5%–10%), and the vast majority of these changes (98%) were associated with genotypes switching from a called to no-call state or vice versa. PCR clones had a larger influence on genotype calls in individuals with low read depth but appeared to influence genotype calls at all loci similarly. Removal of PCR clones reduced the number of called genotypes by 2% but had almost no influence on estimates of heterozygosity. As such, while steps should be taken to limit PCR clones during library preparation, PCR clones are likely not a substantial source of bias for most RAD-seq studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Peter T. Euclide, Garrett McKinney, Matthew Bootsma, Charlene Tarsa, Mariah Meek, Wesley Larson. 2020. Attack of the PCR clones: Rates of clonality have little effect on RAD-seq genotype calls. https://doi.org/10.1111/1755-0998.13087

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

GT-Seq panel development for species identification and parentage analysis of closely related hybridising Scaphirhynchus sturgeons

Hatchery supplementation is vital for conserving dwindling fish populations. Effective augmentation requires distinguishing hatchery-origin from wild individuals and accurately identifying species, particularly in systems where closely related species coexist. Genetic monitoring is key to quantifying genetic differences, but conventional markers do not distinguish hybrids, especially backcrosses. Misidentifying hybrids in hatchery programs compromises wild gene pools because hatchery broodstock contributes to numerous offspring being released into the wild. Here, we present a workflow for developing and evaluating the Genotyping-in-Thousands by sequencing (GT-seq) single nucleotide polymorphism (SNP) panel for North American river sturgeons ( Scaphirhynchus spp.). This panel is designed to detect complex hybrid classes and to determine parent-offspring relationships. Our species identification panel (S-loci) contains 155 SNPs selected for high genetic differentiation (F ST ) between Pallid Sturgeon ( S. albus ) and Shovelnose Sturgeon ( S. platorynchus ), and the parentage assignment panel (P-loci) includes 112 SNPs with high heterozygosity within Pallid Sturgeon. Simulation analyses demonstrated that our GT-seq S-loci panel reliably classifies pure species, F1, F2 and backcross hybrids, even with up to 70% missing data. The P-loci panel achieves high-confidence parentage assignment with ≥ 80% typed loci, with performance influenced by the proportion of sampled parents. Overall, the novel Scaphirhynchus GT-seq panel developed in this study represents a robust and efficient tool for detecting hybridisation, assigning parentage and providing critical information for management decisions in ongoing Pallid Sturgeon conservation.

lower Mississippi River, Missouri River

Development of high-throughput genomic resources to inform white-tailed deer population and disease management

White-tailed deer ( Odocoileus virginianus ) are the most abundant and widespread cervid in North America. Genetic data are used as a tool to monitor populations and make management decisions for this game species. However, the development and use of genomic tools that can generate a set of markers suitable for longitudinal genomic data collection, whether for management purposes or to study the demographic and evolutionary processes of widely distributed species, have been challenging. This is mainly due to the cost required to fully implement and interpret the data produced. Here, we generated whole genome resequencing data for 44 free-ranging deer from three regions in their central and eastern North American range and identified over 89 million single nucleotide polymorphisms (SNPs). We used a subset of these SNPs to develop two nested SNP tools, a high-density array (702,183 SNPs) and a medium-density array (72,723 SNPs) to support deer and chronic wasting disease (CWD) management and research. SNPs were selected to ensure an even distribution across scaffolds of the reference genome and include SNPs associated with CWD susceptibility. Using genotyping results for 469 deer from 15 states in the US and Mexico generated by the high-density array and 1335 deer from 18 states generated by the medium-density array, we assessed genotyping success across different populations and explored some insights into population structure. These genomic tools offer a standard set of markers that will enable researchers and managers to address important questions related to white-tailed deer and CWD management. Our SNP arrays also offer the opportunity to examine aspects of white-tailed deer ecology and evolutionary history that were previously difficult to address.

Molecular Ecology Resources

Sensitive environmental DNA methods for low-risk surveillance of at-risk bumble bees

Terrestrial environmental DNA (eDNA) techniques have been proposed as a means of sensitive, non-lethal pollinator monitoring. To date, however, no studies have provided evidence that eDNA methods can achieve detection sensitivity on par with traditional pollinator surveys. Using a large-scale dataset of eDNA and corresponding net surveys, we show that eDNA methods enable sensitive, species-level characterisation of whole bumble bee communities, including rare and critically endangered species such as the rusty patched bumble bee (RPBB; Bombus affinis ). All species present in netting surveys were detected within eDNA surveys, apart from two rare species in the socially parasitic subgenus Psithyrus (cuckoo bumble bees). Further, for rare non-parasitic species, eDNA methods exhibited similar sensitivity relative to traditional netting. Compared with flower eDNA samples, sequenced leaf surface eDNA samples resulted in significantly lower rates of Bombus detection, and these detections were likely attributable to high rates of background eDNA on environmental surfaces, perhaps due to airborne eDNA or eDNA movement during rainfall events. Lastly, we found that eDNA-based frequency of detection across replicate surveys was strongly associated with net-based measures of abundance across site visits. We conclude that the COI-based metabarcoding method we present is cost-effective and highly scalable for quantitative characterisation of at-risk bumble bee communities, providing a new approach for improving our understanding of species habitat associations.

Central Appalachian Mountains