Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.7.1
Dataset results
219 results for “genotyping‐by‐sequencing”
Data from: Genotyping by sequencing and genome–environment associations in wild common bean predict widespread divergent adaptation to drought
Drought will reduce global crop production by >10% in 2050 substantially worsening global malnutrition. Breeding for resistance to drought will require accessing crop genetic diversity found in the wild accessions from the driest high stress ecosystems. Genome–environment associations in crop wild relatives reveal natural adaptation, and therefore can be used to identify adaptive variation. We explored this approach in the food crop Phaseolus vulgaris L., characterizing 86 geo-referenced wild accessions using Genotyping by Sequencing (GBS) to discover single-nucleotide-polymorphisms (SNPs). The wild beans represented Mesoamerica, Guatemala, Colombia, Ecuador/Northern Peru and Andean groupings. We found high polymorphism with a total of 22,845 SNPs across the 86 accessions loci that confirmed genetic relationships for the groups. As a second objective, we quantified allelic associations with a bioclimatic-based drought index using 10 different statistical models that accounted for population structure. Based on the optimum model, 115 SNPs in 90 regions, widespread in all 11 common bean chromosomes, were associated with the bioclimatic-based drought index. A gene coding for an Ankyrin repeat-containing protein and a phototropic-responsive NPH3 gene were identified as potential candidates. Genomic windows of 1Mb containing associated SNPs had more positive Tajima's D scores than windows without associated markers. This indicates that adaptation to drought, as estimated by bioclimatic variables, has been under natural divergent selection, suggesting that drought tolerance may be favorable under dry conditions but harmful in humid conditions. Our work exemplifies that genomic signatures of adaptation are useful for germplasm characterization, potentially enhancing future marker-assisted selection and crop improvement.
A rapid and versatile tool for HIV-1 Drug Resistance Genotyping by Deep Sequencing: supporting dataset
<p>See file `viral_mixes.md` and [manuscript](http://www.sciencedirect.com/science/article/pii/S0166093416301987).</p> <div class="grammarly-disable-indicator"> </div> <div class="grammarly-disable-indicator"> </div>
Data from: Evaluating genotyping-in-thousands by sequencing as a genetic monitoring tool for a climate sentinel mammal using non-invasive and archival samples
<p>Genetic tools for wildlife monitoring can provide valuable information on spatiotemporal population trends and connectivity, particularly in systems experiencing rapid environmental change. Though many DNA sequencing approaches still require high quality and quantity of DNA obtained from traditional sources (e.g. blood and tissue), rapid genotyping tools such as Genotyping-in-Thousands by sequencing (GT-seq) have improved our ability to make use of degraded and less concentrated DNA commonly obtained from non-invasive and archival samples. Here, we developed a multi-purpose GT-seq panel (307 single nucleotide polymorphisms) for a climate sentinel mammal (the American pika, <em>Ochotona princeps</em>) for use as a genetic tool for monitoring populations in the Canadian Rocky Mountains. We optimized the panel using contemporary tissue samples (n = 77) and subsequently applied it to archival tissue (n = 17) and contemporary fecal pellet samples (n = 129) to evaluate its effectiveness at identifying individuals and sex, estimating relatedness, and inferring population structure. The panel demonstrated high efficacy with contemporary and archival tissue samples (94.7% and 90.5% genotyping success, respectively) and negligible genotyping error (0.001% and 0.0%, respectively). Despite relatively high genotyping success for fecal pellet samples (79.7%), high genotyping error (28.4%) limited its power as a monitoring tool to assess genetic variation using non-invasive samples and highlighted the need for further optimization around sample and data collection.</p>
Genotyping-by-sequencing of Canada's Apple Biodiversity Collection
<p><span>Canada's Apple Biodiversity Collection (ABC) is one of the most diverse collections of apples in the world, which was designed to enable genetic mapping. The ABC is located at the Agriculture and Agri-Food Canada (AAFC) Kentville Research and Development Centre in Nova Scotia, Canada. </span>In addition to phenotypic descriptions of the ABC, sequencing the accessions in the collection provides a valuable resource not only for researchers working on the collection, but for those studying apples more broadly. With this in mind, we report and make publicly available genotyping-by-sequencing (GBS) data for over 1,000 apple accessions from the ABC. By<span> using three SNP callers and imputation, we were able to genotype 278,231 SNPs from 1,175 diverse apple accessions from the ABC.</span></p>
Data from: A genotyping-in-thousands by sequencing panel to inform invasive deer management using non-invasive fecal and hair samples
<p>Studies in ecology, evolution, and conservation often rely on non-invasive samples, making it challenging to generate large amounts of high-quality genetic data for many elusive and at-risk species. We developed and optimized a Genotyping-in-Thousands by sequencing (GT-seq) panel using non-invasive samples to inform the management of invasive Sitka black-tailed deer (<em>Odocoileus hemionus sitkensis</em>) in Haida Gwaii (Canada). We validated our panel using paired high-quality tissue and non-invasive fecal and hair samples to simultaneously distinguish individuals, identify sex and reconstruct kinship among deer sampled across the archipelago, then provided a proof-of-concept application using field-collected feces on SGang Gwaay, an island of high ecological and cultural value. Genotyping success across 244 loci was high (90.3%) and comparable to that of high-quality tissue samples genotyped using restriction-site associated DNA sequencing (92.4%), while genotyping discordance between paired high-quality tissue and non-invasive samples was low (0.50%). The panel will be used to inform future invasive species operations (culls or eradications) in Haida Gwaii by providing individual and population information to inform management. More broadly, our GT-seq workflow that includes quality control analyses for targeted SNP selection and a modified protocol may be of wider utility for other studies and systems where non-invasive genetic sampling is employed.</p>
New methods for the genotyping of Legionella pneumophila - Establishment, validation and implementation of a DNA-based microarray and a core genome multilocus sequence typing
<p>This data presented here are part a doctoral thesis with the focus on new genotyping methods for the human pathogen <em>Legionella pneumophila</em>. The data are partially published in articles. </p> <p>The thesis can be downloaded: update of the URL is coming soon</p>
Genotypes of Aedes aegypti mosquitoes derived from SNP chip and low-coverage whole genome sequencing for platform cross-validation
<p>The mosquito <em>Aedes aegypti </em>is the primary vector of many human arboviruses such as dengue, yellow fever, chikungunya, and Zika, which affect millions of people world-wide. Population genetics studies on this mosquito have been important in understanding its invasion pathways and success as a vector of human disease. The Axiom aegypti1 SNP chip was developed from a sample of geographically diverse <em>Ae. aegypti </em>populations to facilitate genomic studies on this species. Here we evaluate the utility of the Axiom aegypti1 SNP chip for population genetics and compare it with a low-depth shot-gun sequencing approach using mosquitoes from the species' native (Africa) and invasive range (outside Africa). These analyses indicate that the results from the SNP chip are highly reproducible and have a higher sensitivity to capture alternative alleles than a low-coverage whole-genome sequencing approach. Although the SNP chip suffers from ascertainment bias, results from population structure, ancestry, demographic, and phylogenetic analyses using the SNP chip were congruent with those derived from low coverage whole genome sequencing, and consistent with previous reports on Africa and outside Africa populations using microsatellites. More importantly, we identified a subset of SNPs that can be reliably used to generate merged databases, opening the door to combined analyses. We conclude that the Axiom aegypti1 SNP chip is a convenient, more accurate, low-cost alternative to low-depth whole genome sequencing for population genetic studies of <em>Ae. aegypti</em> that do not rely on full allelic frequency spectra. Whole genome sequencing and SNP chip data can be easily merged, extending the usefulness of both approaches. </p>
Data from: Targeted genotyping-by-sequencing of potato and data analysis with R/polyBreedR
<p>"Mid-density" targeted genotyping-by-sequencing (GBS) combines trait-specific markers with thousands of genomic markers at an attractive price for linkage mapping and genomic selection. A 2.5K targeted GBS assay for potato was developed using the DArTag<sup>TM</sup> technology and later expanded to 4K targets. Genomic markers were selected from the potato Infinium<sup>TM</sup> SNP array to maximize genome coverage and polymorphism rates. The DArTag and SNP array platforms produced equivalent dendrograms in a test set of 298 tetraploid samples, and 83% of the common markers showed good quantitative agreement, with RMSE (root-mean-squared-error) less than 0.5. DArTag is suited for genomic selection candidates in the clonal evaluation trial, coupled with imputation to a higher-density platform for the training population. Using the software polyBreedR, an R package for the manipulation and analysis of polyploid marker data, the RMSE for imputation by linkage analysis was 0.15 in a small half-diallel population (N=85), which was significantly lower than the RMSE of 0.42 with the Random Forest method. Regarding high-value traits, the DArTag markers for resistance to potato virus Y, golden cyst nematode, and potato wart appeared to track their targets successfully, as did multi-allelic markers for maturity and tuber shape. In summary, the potato DArTag assay is a transformative and publicly available technology for potato breeding and genetics.</p>
Distribution of Chlamydia trachomatis ompA-genotypes over three decades (1990-2021) in Portugal - ompA sequence datasets
<p>This repository includes sequence data from the study: "<strong>Distribution of <em>Chlamydia trachomatis</em> <em>ompA</em>-genotypes over three decades (1990-2021) in Portugal</strong>”, conducted by the <strong>National Reference Laboratory (NRL) for Sexually Transmitted Infections (STI), National Institute of Health Doutor Ricardo Jorge (INSA), Portugal</strong>.</p> <p>The NRL performs molecular characterization (namely <em>ompA</em>-genotyping) of all <em>C. trachomatis</em> positive samples that receives. <em>C. trachomatis</em> <em>ompA</em>-genotyping technique was adapted from Lan, et al. [1]. Briefly, PCR and nested PCR were performed using primers NLO and NRO, and primers PCTM3 and SERO2A, respectively, as previously described [2]. Partial nucleotide sequencing of the ~1010 bp PCR product was performed with BigDye terminator v1.1 and capillary sequencing (3130XL Genetic Analyzer; Applied Biosystems), using either two primers, as described elsewhere [2,3,4], or one primer (since ~2018), as described elsewhere [5]. Until 2018, LaserGene (DNASTAR) and MEGA (http://www.megasoftware.net) software were applied for sequence curation, alignment and phylogenetic reconstructions involving multiple reference sequences, as previously described [3]. Since ~2018, <em>ompA</em> genotypes have been determined by BLASTn-based comparison (using the ABRIcate tool [6]) directly from raw Sanger (AB1 format) sequences, with a custom database enrolling reference and variant sequences of all main <em>ompA </em>genotypes (<a href="https://github.com/insapathogenomics/ReporType/blob/main/databases/c_trachomatis.fasta">https://github.com/insapathogenomics/ReporType/blob/main/databases/c_trachomatis.fasta</a>) [3, 5, 7, 8], as currently implemented in <strong>ReporType</strong> (<a href="https://github.com/insapathogenomics/ReporType">https://github.com/insapathogenomics/ReporType</a>) [8]. When needed, MEGA software is then applied for fine genotype confirmation, namely for L2/L2b discrimination and confirmation of the hybrid <em>ompA</em>-profile of the recombinant L2b/D-Da [5].</p> <p>This repository includes the following sequence datasets:</p> <ul> <li><strong>Dataset 1</strong> - <em>ompA</em> sequences (curated FASTA) of <em>C. trachomatis</em> positive samples collected between 1991 and ~2018, as described above.</li> <li><strong>Dataset 2</strong> - <em>ompA</em> sequences (raw Sanger sequences, converted from “ab1” format to FASTA) of <em>C. trachomatis</em> positive samples collected since ~2018 and 2021, as described above.</li> </ul> <p><em>Note: The associated metadata is described in the Supplementary table 2 of the manuscript. These sequence datasets do not cover genotyped samples for which the ompA sequences were lost over the three decades of the laboratory's historical collection.</em></p> <p> </p> <p>References</p> <p>1. Lan J, Ossewaarde JM, Walboomers JM, Meijer CJ, van den Brule AJ. Improved PCR sensitivity for direct genotyping of Chlamydia trachomatis serovars by using a nested PCR. <em>J Clin Microbiol</em>. 1994;32(2):528-530. doi:10.1128/jcm.32.2.528-530.1994;</p> <p>2. Gomes JP, Bruno WJ, Borrego MJ, Dean D. Recombination in the genome of <em>Chlamydia trachomatis </em>involving the polymorphic membrane protein C gene relative to <em>ompA </em>and evidence for horizontal gene transfer. <em>J Bacteriol </em>2004;186:4295–4306;</p> <p>3. Nunes A, Borrego MJ, Nunes B, Florindo C, Gomes JP. Evolutionary dynamics of <em>ompA</em>, the gene encoding the <em>Chlamydia trachomatis</em> key antigen. J Bacteriol. 2009 Dec;191(23):7182-92. doi: 10.1128/JB.00895-09. Epub 2009 Sep 25. PMID: 19783629; PMCID: PMC2786549;</p> <p>4. Nunes A, Nogueira PJ, Borrego MJ, Gomes JP. Adaptive evolution of the Chlamydia trachomatis dominant antigen reveals distinct evolutionary scenarios for B- and T-cell epitopes: worldwide survey. PLoS One. 2010 Oct 5;5(10):e13171. doi: 10.1371/journal.pone.0013171. PMID: 20957150; PMCID: PMC2950151.</p> <p>5. Borges V, Cordeiro D, Salas AI, et al. <em>Chlamydia trachomatis</em>: when the virulence-associated genome backbone imports a prevalence-associated major antigen signature. <em>Microb Genom</em>. 2019;5(11):e000313. doi:10.1099/mgen.0.000313;</p> <p>6. Seemann T. ABRIcate. <a href="https://github.com/tseemann/abricate">https://github.com/tseemann/abricate</a></p> <p>7 Nunes A, Nogueira PJ, Borrego MJ, Gomes JP. Adaptive evolution of the Chlamydia trachomatis dominant antigen reveals distinct evolutionary scenarios for B- and T-cell epitopes: worldwide survey. PLoS One. 2010 Oct 5;5(10):e13171. doi: 10.1371/journal.pone.0013171. PMID: 20957150; PMCID: PMC2950151.</p> <p>8. Cruz H, Pinheiro M, Borges V. ReporType: a flexible bioinformatics tool for targeted loci screening and typing of infectious agents (<a href="https://github.com/insapathogenomics/ReporType">https://github.com/insapathogenomics/ReporType</a>). Int J Mol Sci. 2024;25:3172. https://doi.org/10.3390/ijms25063172</p>
Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference
Restriction site-associated DNA sequencing (RADseq) provides researchers with the ability to record genetic polymorphism across thousands of loci for non-model organisms, potentially revolutionising the field of molecular ecology. However, as with other genotyping methods, RADseq is prone to a number of sources of error that may have consequential effects for population genetic inferences, and these have received only limited attention in terms of the estimation and reporting of genotyping error rates. Here we use individual sample replicates, under the expectation of identical genotypes, to quantify genotyping error in the absence of a reference genome. We then use sample replicates to (1) optimize de novo assembly parameters within the program Stacks, by minimizing error and maximizing the retrieval of informative loci, and; (2) quantify error rates for loci, alleles and SNPs. As an empirical example we use a double digest RAD dataset of a non-model plant species, Berberis alpina, collected from high altitude mountains in Mexico.
Data from: A RAD-sequencing approach to genome-wide marker discovery, genotyping, and phylogenetic inference in a diverse radiation of primates
Until recently, most phylogenetic and population genetics studies of nonhuman primates have relied on mitochondrial DNA and/or a small number of nuclear DNA markers, which can limit our understanding of primate evolutionary and population history. Here, we describe a cost-effective reduced representation method (ddRAD-seq) for identifying and genotyping large numbers of SNP loci for taxa from across the New World monkeys, a diverse radiation of primates that shared a common ancestor ~20-26 mya. We also estimate, for the first time, the phylogenetic relationships among 15 of the 22 currently-recognized genera of New World monkeys using ddRAD-seq SNP data using both maximum likelihood and quartet-based coalescent methods. Our phylogenetic analyses robustly reconstructed three monophyletic clades corresponding to the three families of extant platyrrhines (Atelidae, Pitheciidae and Cebidae), with Pitheciidae as basal within the radiation. At the genus level, our results conformed well with previous phylogenetic studies and provide additional information relevant to the problematic position of the owl monkey (Aotus) within the family Cebidae, suggesting a need for further exploration of incomplete lineage sorting and other explanations for phylogenetic discordance, including introgression. Our study additionally provides one of the first applications of next-generation sequencing methods to the inference of phylogenetic history across an old, diverse radiation of mammals and highlights the broad promise and utility of ddRAD-seq data for molecular primatology.
Alfalfa genotyping-by-sequencing (GBS) data
<p>Alfalfa (<i>Medicago</i> <i>sativa</i> L.) quantitative trait loci (QTL) mapping population (184 F<sub>1</sub>) derived from cultivars 3010 (cold-tolerant) as female parent and CW 100 (cold-sensitive) as male parent were genotyped using genotyping-by-sequencing (GBS). Polymorphic SNPs unique to either 3010 (AB x AA) or CW 1010 (AA x AB) were identified as single dose allele (SDA) markers and used to generate the genetic linkage maps. Two sets of linkage maps, a set for each parent, were used to map the traits and the QTL were identified. With the genotyping and phenotyping informations we were able to map various alfalfa traits such as fall dormancy, winter-hardiness, freezing tolerance, flowering time, yield and leaf-rust resistance. The raw sequence data were deposited at NCBI SRA with the accession number SRP150116. This study identified several genomic regions and associated markers that can be further utilized in marker-assisted breeding to improve the alfalfa. </p>
Improved library preparation protocols for amplicon sequencing-based noninvasive fetal genotyping for RHD-positive D antigen-negative alleles
<p>We aimed to simplify our fetal <i>RHD</i> genotyping protocol by changing the method to attach Illumina's sequencing adaptors to PCR products from the ligation-based method to a PCR-based method, and to improve its quantitative accuracy by introducing unique molecular indexes, which allow us to count the numbers of DNA fragments used as PCR templates and to minimize the effects of PCR and sequencing errors. Both of the newly established protocols reduced time and cost compared with our conventional protocol. Removal of PCR duplicates using UMIs reduced the frequencies of erroneously mapped sequences reads likely generated by PCR and sequencing errors. The modified protocols will help us facilitate implementing fetal <i>RHD</i> genotyping for East Asian populations into clinical practice.</p>
Data from: A comparison of non-destructive visceral swab and tissue biopsy sampling methods for genotyping-by-sequencing in the freshwater mussel Fusconaia askewi
<p>Limiting harm to organisms via genetic sampling is an important consideration for rare species. Nondestructive sampling techniques have been developed to address this issue in freshwater mussels. Two methods, visceral swabbing and tissue biopsies, have proven to be effective for DNA sampling, though it is unclear as to which method is preferable for genotyping-by-sequencing (GBS). Tissue biopsies may cause undue stress and damage to organisms, while visceral swabbing potentially reduces the chance of such harm. Our study compared the efficacy of these two DNA sampling methods for generating GBS data for the Unionid freshwater mussel, Texas Pigtoe (<em>Fusconaia askewi</em>). Our results find both methods generate quality sequence data, though some considerations are in order. Tissue biopsies produced significantly higher DNA concentrations and larger numbers of reads when compared to swabs, though there was no significant association between starting DNA concentration and number of reads generated. Swabbing produced greater sequence depth (more reads per sequence) while tissue biopsies revealed greater coverage across the genome (at lower sequence depth). Patterns of genomic variation as characterized in principal component analyses were similar regardless of the sampling method, suggesting that the less invasive swabbing is a viable option for producing quality GBS data in these organisms.</p>
Alfalfa genotyping-by-sequencing (GBS) data
Open the record for dataset details and reuse information.
Colonization history of the Canary Islands endemic Lavatera acerifolia, (Malvaceae) unveiled with Genotyping-by-Sequencing data and niche modeling
Open the record for dataset details and reuse information.
Improved library preparation protocols for amplicon sequencing-based noninvasive fetal genotyping for RHD-positive D antigen-negative alleles
Open the record for dataset details and reuse information.
Development and application of Faba_bean_130K Targeted Next-Generation Sequencing SNP genotyping platform based on transcriptome sequencing
Open the record for dataset details and reuse information.
GBS SNP datasets from "Genotyping-by-sequencing resolves relationships in Polygonaceae tribe Eriogoneae", TAXON
Open the record for dataset details and reuse information.
Data from: A RAD-sequencing approach to genome-wide marker discovery, genotyping, and phylogenetic inference in a diverse radiation of primates
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.