Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
549
datasets available to search
ShareScore release 0.9.0
Dataset results
549 results for “SNP data”
Western redcedar single nucleotide polymorphism (SNP) genotyping data for genomic selection and population genetics
<p>Western redcedar (<em>Thuja plicata</em>) Single Nucleotide Polymorphism (SNP) data in Variant Call Format (VCF) for genomic selection training and target populations, genomic selection parents, and self-fertilized (selfing) lines, comprising 4,833 trees.</p> <p>Targeted sequencing-based genotyping was done by Capture-Seq methodology at Rapid Genomics (Neves est al. 2013). A set of 57,000 probes as designed for initial marker discovery, from which a panel of 20,858 probes was selected for genotyping. A set of transcriptomes (Shalev et al. 2018) (PRJNA704616) was aligned to the reference genome to identify SNPs. Candidate probes (120 nt) were initially designed in silico and 57,000 selected by removing candidates with poor base composition for hybridization (GC content <0.2 and >0.6, high G content >0.2 and long homopolymers >7), followed by removing probes aligning to more than one position on the reference genome (≥90% identity and length). The 57,000 probes represent 14,517 scaffolds (average 3.9 probes/scaffold), with 37,275 targeting at least one SNP and 19,725 mapping to intergenic regions not containing pre-identified SNPs. A set of 128 individuals were selected to validate the 57,000 probe panel and associated polymorphisms. Genomic DNA (0.5 ug) was fragmented (mean size 300 bp), followed by repair of ends, phosphorylation, adenylation, ligation of Illumina compatible adapters containing 8bp indexes and 5’ T-overhang, and 10 cycles PCR amplification with universal primers to produce sequencing-ready libraries. Libraries were quantified using PicoGreen. Libraries from 16 samples were pooled, hybridized to the 120 nt RNA probes following Agilent’s SureSelect Target Enrichment System (Agilent Technologies) and sequenced on an Illumina HiSeq X machine with paired-end 150bp cycle for an average sequencing depth per sample of 15X. Sequence data were aligned to the reference genome with BWA-MEM (http://arxiv.org/abs/1303.3997) and sets of four samples were combined to increase sequencing depth for identifying markers. Putative SNPs were identified using Freebayes (http://arxiv.org/abs/1207.3907) in 150bp on either side of the 57,000 probes and filtered probes that had more than 17 SNPs per 420 bp target region (150bp + 120bp + 150bp). The sequencing depth of the probes was used to select the final set of 20,885 probes, removing probes on both sides of the distribution (low and high sequencing depth), for Capture-Seq on the remainder of the samples.</p>
Sorghum bicolor SNP data
<div> <div> <p>Sorghum diversity set was utilized for the present study.</p> </div> </div> <div> </div>
Data from: "Identification of SNP markers for the endangered Ugandan red colobus (Procolobus rufomitratus tephrosceles) using RAD sequencing" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015
Despite dramatic growth in the field of primate genomics over the past decade, studies of primate population and conservation genomics in the wild have been hampered due to the difficulties inherent in studying non-model organisms and endangered species, such as lack of a reference genome and current challenges in de novo primate genome assembly. Here, we used Restriction-site Associated DNA (RAD) sequencing to develop a population-based SNP panel for the Ugandan red colobus (P. rufomitratus tephrosceles), which is a highly threatened monkey due to habitat loss. We analyzed blood samples from 24 individuals from Kibale National Park (Uganda) using single-end RAD sequencing. We obtained 70,773,857 reads, of which 58,814,906 passed the filtering steps. Using the program STACKS v. 1.11 we identified 113,376 loci, of which 50,558 were polymorphic and had a mean observed heterozygosity of 0.25. These data will be used to study the effects of habitat fragmentation on genomic diversity, dispersal, and disease transmission in this species. Our approach provides a good example of the potential of RAD sequencing in studies of wild primate populations.
Data from: Number of alleles as a predictor of the relative assignment accuracy of STR and SNP baselines for chum salmon
Short tandem repeat (STR) markers, which exhibit many alleles per locus, are commonly used to assign fish to their populations of origin. Single nucleotide polymorphisms (SNPs), which have many technical advantages over STRs, typically exhibit only two alleles per locus. Simulation studies have indicated that number of independent alleles is a good predictor of accuracy of genetic markers for fishery applications. Extant STR baselines for salmon contain hundreds of alleles, and it has been extrapolated that hundreds of SNP markers need to be developed before SNP baselines will compare to these STR baselines. We compared 15 STRs exhibiting 349 independent alleles to 61 SNP assays exhibiting 66 independent alleles for accuracy in assigning to closely related populations of chum salmon. The SNP baseline yielded slightly higher mean accuracies for proportional assignment and comparable accuracies for individual assignment. Overall the SNP baseline performed considerably better, relative to the microsatellite baseline, than predicted based on the number of independent alleles in each baseline. We suggest that this discrepancy is due to the fact that the simulation studies do not capture the impacts of the different strategies commonly employed for discovering and selecting STR and SNP markers.
Data from: Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks
An Illumina Infinium SNP genotyping array was constructed for European white oaks. Six individuals of Quercus petraea and Q. robur were considered for SNP discovery using both previously obtained Sanger sequences across 676 gene regions (1371 in vitro SNPs) and Roche 454 technology sequences from 5112 contigs (6542 putative in silico SNPs). The 7913 SNPs were genotyped across the six parental individuals, full-sib progenies (one within each species and two interspecific crosses between Q. petraea and Q. robur) and three natural populations from south-western France that included two additional interfertile white oak species (Q. pubescens and Q. pyrenaica). The genotyping success rate in mapping populations was 80.4% overall and 72.4% for polymorphic SNPs. In natural populations, these figures were lower (54.8% and 51.9%, respectively). Illumina genotype clusters with compression (shift of clusters on the normalized x-axis) were detected in ~25% of the successfully genotyped SNPs and may be due to the presence of paralogues. Compressed clusters were significantly more frequent for SNPs showing a priori incorrect Illumina genotypes, suggesting that they should be considered with caution or discarded. Altogether, these results show a high experimental error rate for the Infinium array (between 15% and 20% of SNPs potentially unreliable and 10% when excluding all compressed clusters), and recommendations are proposed when applying this type of high-throughput technique. Finally, results on diversity levels and shared polymorphisms across targeted white oaks and more distant species of the Quercus genus are discussed, and perspectives for future comparative studies are proposed.
Data from: Genetic diversity, linkage disequilibrium and selection signatures in Chinese and Western pigs revealed by genome-wide SNP markers
To investigate population structure, linkage disequilibrium (LD) pattern and selection signature at the genome level in Chinese and Western pigs, we genotyped 304 unrelated animals from 18 diverse populations using porcine 60 K SNP chips. We confirmed the divergent evolution between Chinese and Western pigs and showed distinct topological structures of the tested populations. We acquired the evidence for the introgression of Western pigs into two Chinese pig breeds. Analysis of runs of homozygosity revealed that historical inbreeding reduced genetic variability in several Chinese breeds. We found that intrapopulation LD extents are roughly comparable between Chinese and Western pigs. However, interpopulation LD is much longer in Western pigs compared with Chinese pigs with average r20.3 values of 125 kb for Western pigs and only 10.5 kb for Chinese pigs. The finding indicates that higher-density markers are required to capture LD with causal variants in genome-wide association studies and genomic selection on Chinese pigs. Further, we looked across the genome to identify candidate loci under selection using FST outlier tests on two contrast samples: Tibetan pigs versus lowland pigs and belted pigs against non-belted pigs. Interestingly, we highlighted several genes including ADAMTS12, SIM1 and NOS1 that show signatures of natural selection in Tibetan pigs and are likely important for genetic adaptation to high altitude. Comparison of our findings with previous reports indicates that the underlying genetic basis for high-altitude adaptation in Tibetan pigs, Tibetan peoples and yaks is likely distinct from one another. Moreover, we identified the strongest signal of directional selection at the EDNRB loci in Chinese belted pigs, supporting EDNRB as a promising candidate gene for the white belt coat color in Chinese pigs. Altogether, our findings advance the understanding of the genome biology of Chinese and Western pigs.
Data from: SNP discovery in European lobster (Homarus gammarus) using RAD sequencing
The European lobster (Homarus gammarus) is a decapod crustacean with a high market value and therefore their fisheries are of major importance to the economies they support. However, over-exploitation has led to profound stock declines in some regions such as Scandinavia and the Mediterranean. To manage this resource sustainably, knowledge of population structure and connectivity is crucial to inform management about dispersal, recruitment, stock identification and food traceability. We used restriction-site associated DNA sequencing to develop novel SNP markers from 55 individuals encompassing much of the species range; SNPs were quality filtered, ranked using F-statistics and the top 96 SNPs adequate for primer design were retained. SNP markers were developed with the aim of maximising the power to detect genetic differentiation between: (i) Atlantic and Mediterranean lobsters and (ii) Atlantic lobsters. This panel of SNPs provides a useful resource for future studies of population genetic structure and assignment in H. gammarus.
Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements
Sample availability limits population genetics research on many species, especially taxa from regions with high diversity. However, many such species are well represented in museum collections assembled before the molecular era. Development of techniques to recover genetic data from these invaluable specimens will benefit biodiversity science. Using a mixture of freshly preserved and historical tissue samples, and a sequence capture probe set targeting >5000 loci, we produced high-confidence genotype calls on thousands of single nucleotide polymorphisms (SNPs) in each of five South-East Asian bird species and their close relatives (N = 27–43). On average, 66.2% of the reads mapped to the pseudo-reference genome of each species. Of these mapped reads, an average of 52.7% was identified as PCR or optical duplicates. We achieved deeper effective sequencing for historical samples (122.7×) compared to modern samples (23.5×). The number of nucleotide sites with at least 8× sequencing depth was high, with averages ranging from 0.89 × 106 bp (Arachnothera, modern samples) to 1.98 × 106 bp (Stachyris, modern samples). Linear regression revealed that the amount of sequence data obtained from each historical sample (represented by per cent of the pseudo-reference genome recovered with ≥8× sequencing depth) was positively and significantly (P ≤ 0.013) related to how recently the sample was collected. We observed characteristic post-mortem damage in the DNA of historical samples. However, we were able to reduce the error rate significantly by truncating ends of reads during read mapping (local alignment) and conducting stringent SNP and genotype filtering.
Mimosa catherinensis SNP data sets
<p>To inform management strategies for conservation of <i>Mimosa catharinensis </i>– a narrow endemic, critically endangered plant species – we identified 1,497 unlinked SNP markers derived from a reduced representation sequencing method (i.e., ddRADseq). This set of molecular markers was employed to assess intrapopulation genetic parameters and the demographic history of one extremely small population of <i>M. catharinensis </i>located in the Brazilian Atlantic Forest. We observed a moderate level of genetic diversity for <i>M. catharinensis</i>. Interestingly, <i>M. catharinensis</i>, which is a lianescent shrub with no indication of seed production for at least two decades, presented high levels of outcrossing and no evidence of inbreeding. However, the reconstruction of demographic history of <i>M. catharinensis</i> indicate that the population should be suffered a recent bottleneck.</p>
Genetic loci associated with winter survivorship in diverse lowland switchgrass populations: SNP read count data
<p>High winter mortality is the most important factor limiting biomass yield of lowland switchgrass planted in the northern latitudes of North America. Due to the perennial growth habit and strong dependence on weather conditions to generate sufficient selection pressure to identify winter-hardy individuals, breeding of cold tolerant switchgrass cultivars requires many years. Identification of causal genetic variants for winter survivorship would accelerate the improvement of switchgrass biomass production. The objective of this study was to identify allelic variation associated with winter survivorship in lowland switchgrass populations using bulk segregant analysis (BSA). Twenty-nine lowland switchgrass populations were evaluated for winter survival at two locations in southern Wisconsin and 21 population with differential winter survivorship was used for BSA. A maximum of 10% of the individuals per population (8-20) was bulked to create survivor and non-survivor DNA pools. The DNA pools were evaluated using exome capture sequencing and allele frequencies were used to conduct statistical tests. The BSA tests revealed nine QTL from tetraploid populations and seven QTL from octoploid populations. Some markers were identified in multiple populations that originated across a broad geographic landscape, while other markers were site-specific. QTL at positions 88 Mb on chromosome 2N, 115 Mb on chromosome 5K, and 1 and 100 Mb on chromosome 9N were potentially the most useful QTL. Markers associated with winter survivorship in this study can be used to accelerate breeding cycles of lowland switchgrass populations and should lead to improvements in adaptation within USDA hardiness zones 4 and 5.</p>
SNP data set of the Peruvian Creole cattle from southern Peru
<p>The Peruvian creole cattle (PCC) was originated after the introduction of cattle into the American continent about five centuries ago, and is an important source of power for agriculture, meat, and milk in the Peruvian highlands, as well as part of cultural traditions. However, little is known about the genetics of the PCC. In order to determine the genetic diversity and structure of the PCC, 69 DNA samples from four southern regions of Peru (Apurimac, Ayacucho, Cusco and Puno) were genotyped using a 100K SNP bead chip. After quality control and LD pruning, 24,200 SNPs were retained for further analysis. Animals were grouped into two clusters (C1: Apurimac, Ayacucho and Cusco, C2: Puno) using principal component analysis and UPGMA dendrogram. STRUCTURE analysis showed that individuals from Puno grouped in one cluster. Expected heterozygosity ranged from 0.399 (Apurimac) to 0.418 (Ayacucho). Negative inbreeding coefficient (F<sub>IS</sub>) values for PCC from Puno and Ayacucho were also found, possibly due to admixture. The lowest F<sub>ST</sub> (0.005) was estimated for Ayacucho and Cusco cattle populations, and the highest F<sub>ST</sub> (0.028) was reported for Puno and Apurimac cattle population. Small genetic variation among populations (3.65%) but higher variation within populations was found using AMOVA. To the best of our knowledge, this is the first study employing SNP markers in PCC, and as such it is hoped that this helps to pave the way towards its genetic improvement and the urgent sustainable management of creole animals in Peru.</p>
Data from: Estimating genomic diversity and population differentiation – an empirical comparison of microsatellite and SNP variation in Arabidopsis halleri
Open the record for dataset details and reuse information.
Data from: Accuracy of assignment of Atlantic salmon (Salmo salar L.) to rivers and regions in Scotland and northeast England based on single nucleotide polymorphism (SNP) markers.
Open the record for dataset details and reuse information.
Data from: Distinguishing the victim from the threat: SNP‐based methods reveal the extent of introgressive hybridization between wildcats and domestic cats in Scotland and inform future in situ and ex situ management options for species restoration
Open the record for dataset details and reuse information.
Data from: A comparative assessment of SNP and microsatellite markers for assigning parentage in a socially monogamous bird
Open the record for dataset details and reuse information.
Data from: A study of applicability of SNP chips developed for bovine and ovine species to whole-genome analysis of reindeer Rangifer tarandus
Open the record for dataset details and reuse information.
Data from: A cost-and-time effective procedure to develop SNP markers for multiple species: a support for community genetics
Open the record for dataset details and reuse information.
Archival and modern DNA SNP data of 13 Baltic salmon populations
Open the record for dataset details and reuse information.
Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements
Open the record for dataset details and reuse information.
Data from: Insights into the genetic history of French cattle from dense SNP data on 47 worldwide breeds
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.