Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
549
datasets available to search
ShareScore release 0.9.0
Dataset results
549 results for “SNP data”
Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs
<p>The development of high-throughput sequencing has prompted a transition in wildlife genetics from using microsatellites toward sets of Single Nucleotide Polymorphisms (SNPs). However, genotyping large numbers of targeted SNPs using non-invasive samples remains challenging due to relatively large DNA input requirements. Recently, target enrichment has emerged as a promising approach requiring little template DNA. We assessed the efficacy of Tecan Genomics' Allegro Targeted Genotyping (ATG) for generating genome-wide SNP data in feral horses using DNA isolated from fecal swabs. Total and host-specific DNA were quantified for 989 samples collected as part of a long-term individual-based study of feral horses on Sable Island, Nova Scotia, Canada, using dsDNA fluorescence and a host-specific qPCR assay, respectively. Forty-eight samples representing 44 individuals containing at least 10ng of host DNA (ATG's recommended minimum input) were genotyped using a custom multiplex panel targeting 279 SNPs. Genotyping accuracy and consistency were assessed by contrasting ATG genotypes with those obtained from the same individuals with SNP microarrays, and from multiple samples from the same horse, respectively. 62% of swabs yielded the minimum recommended amount of host DNA for ATG. Ignoring samples that failed to amplify, ATG recovered an average of 86.7% targeted sites per sample, while genotype concordance between ATG and SNP microarrays was 98.5%. The repeatability of genotypes from the same individual approached unity with an average of 99.9%. This study demonstrates the suitability of ATG for genome-wide, non-invasive targeted SNP genotyping, and will facilitate further ecological and conservation genetics research in equids and related species.</p>
Dartseq SNP data of Uganda sorghum germplasm
<p>SNP data generated DArTseq for Ugandan S. bicolor germplasm accessions (UG set) from the Plant Genetic Resources Centre at the Uganda National Genebank.</p>
Data from high throughput SNP-chip as cost effective new monitoring tool for assessing invasion dynamics in the comb jelly Mnemiopsis leidyi
<p class="MsoNormal"><span>High throughput low-density SNP arrays provide a cost-effective solution for population genetic studies and monitoring of genetic diversity as well as population structure commonly implemented in real time stock assessment of fish species. However, the application of high throughput SNP arrays for monitoring of invasive species has so far not been implemented. We developed a species-specific SNP array for the invasive comb jelly <em>Mnemiopsis leidyi</em> based on whole genome resequencing data. Initially, </span><span>a total of</span><span> </span><span>1,</span><span>395</span><span> </span><span>high quality </span><span>SNPs</span><span> were identified</span><span> </span><span>u</span><span>sing stri</span><span>ngent</span><span> filtering criteria</span><span>. From those, 192 assays were designed and validated, resulting in the final panel of 116 SNPs. Markers were diagnostic between the northern and southern <em>M. leidyi</em> lineages and highly polymorphic to distinguish populations. Despite using a reduced representation of the genome, our SNP panel yielded comparable results to using a whole genome resequencing approach (832,323 SNPs), recovering similar values of genetic differentiation between samples and detecting the same clustering groups when performing Structure analyses. The resource presented here provides a cost-effective, high throughput solution for population genetic studies, allowing to routinely genotype large number of individuals. Monitoring of genetic diversity and effective population size estimations in this highly invasive species will allow for the early detection of new introductions from distant source regions or hybridization events. Thereby, this SNP chip represents an important management tool in order to understand invasion dynamics and </span><span>opens the door for implementing such methods for a wider range of alien invasive species.</span></p> <div></div>
SNP data (DArTseq) for population genomics of Araucaria bidwillii
<p><span>We took Araucaria bidwillii leaf DNA samples from a total of 31 sites and 171 samples, representing 3 sites from a northern population in the Australian Wet Tropics and 28 sites from a southern population in Southeast Queensland, Australia. </span>SNP data was obtained from genotyping-by-sequencing platform Divesity Arrays Technology (DArTseq) and the resultant dataset has not been processed for quality control.</p>
Autosomal SNP-genotype data of brown bears (Ursus arctos) in Finland
<p>Harmonising methodology between countries is crucial in transborder population monitoring. However, immediate application of alleged, established DNA-based methods across the extended area can entail drawbacks and may lead to biases. Therefore, genetic methods need to be tested across the whole area before being deployed. Around 4,500 brown bears (<em>Ursus arctos</em>) live in Norway, Sweden, and Finland and they are divided into the western (Scandinavian) and eastern (Karelian) population. Both populations have recovered and are connected via asymmetric migration. DNA-based population monitoring in Norway and Sweden uses the same set of genetic markers. With Finland aiming to implement monitoring, we tested the available SNP-panel developed to assess brown bears in Norway and Sweden, on tissue samples from a representative set of 93 legally harvested individuals from Finland. The aim was to test for ascertainment bias and evaluate its suitability for DNA-based transnational-monitoring covering all three countries. We compared results to the performance of microsatellite genotypes of the same individuals in Finland and against SNP-genotypes from individuals sampled in Sweden (<em>N</em>=95) and Norway (<em>N</em>=27). In Finland, a higher resolution for individual identification was obtained for SNPs (PI=1.18E-27) compared to microsatellites (PI=4.2E-11). Compared to Norway and Sweden, probability of identity of the SNP-panel was slightly higher and expected heterozygosity lower in Finland indicating ascertainment bias. Yet, our evaluation show that the available SNP-panel outperforms the microsatellite panel currently applied in Norway and Sweden. The SNP-panel represents a powerful tool that could aid improving transnational DNA-based monitoring of brown bears across these three countries.</p>
SNP data set of Peruvian highland maize races
<p>Peruvian maize exhibits significant morphological diversity, with landraces cultivated from sea level up to 3,500 meters above sea level. Previous research based on morphological descriptors identified at least 52 Peruvian maize races, but their genetic diversity and population structure remain largely unknown. In this study, we used genotyping-by-sequencing (GBS) to infer the genetic structure and diversity of 423 maize accessions from the Genebank of La Molina National Agrarian University (UNALM). These accessions represent nine races and one sub-race, along with 15 open-pollinated lines (purple corn) and two yellow maize hybrids. We obtained 14,235 high-quality SNPs distributed along the 10 maize chromosomes. Gene diversity ranged from 0.33 (Pachia) to 0.362 (Ancashino), with Cusco showing the lowest inbreeding coefficient (0.205) and Ancashino the highest (0.274) among the landraces. Population divergence (FST) was very low (mean = 0.017), indicating extensive interbreeding among Peruvian maize varieties. Population structure analysis revealed that these 423 distinct genotypes could be grouped into 10 clusters, with some maize races clustering together. Peruvian maize races did not form monophyletic groups; instead, our phylogenetic tree identified two clades corresponding to the chronological classification of Peruvian maize races: <em>Anciently Derived or Primary Races</em> (ADPR) and <em>Lately Derived or Secondary Races</em> (LDSR). These clades also align with the geographic origins of the maize races, reflecting their mixed evolutionary backgrounds. Further investigation of Peruvian maize germplasm using modern technologies is essential to enhance their use in breeding programs, particularly in the Andean region of Peru.</p>
Concatenated SNP data: Integrative taxonomy of the lizards Cercosaura ocellata species complex (Reptilia: Gymnophthalmidae) based on morphological and genomic data
<p>Concatenated unliked SNP data in phylip format used in 'Integrative taxonomy of the lizards Cercosaura ocellata species complex (Reptilia: Gymnophthalmidae) based on morphological and genomic data' study.</p>
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
<p>The profiling of epigenetic marks like DNA methylation has become a central aspect of studies in evolution and ecology. Bisulfite sequencing is commonly used for assessing genome-wide DNA methylation at single nucleotide resolution but these data can also provide information on genetic variants like single nucleotide polymorphisms (SNPs). However, bisulfite conversion causes unmethylated cytosines to appear as thymines, complicating the alignment and subsequent SNP calling. Several tools have been developed to overcome this challenge, but there is no independent evaluation of such tools for non-model species, which often lack genomic references. Here, we used whole-genome bisulfite sequencing (WGBS) data from four female great tits (<i>Parus major</i>) to evaluate the performance of seven tools for SNP calling from bisulfite sequencing data. We used SNPs from whole-genome resequencing data of the same samples as baseline SNPs to assess common performance metrics like sensitivity, precision, and the number of true positive, false positive, and false negative SNPs for the full range of variant and genotype quality values. We found clear differences between the tools in either optimizing precision (Bis-SNP), sensitivity (biscuit), or a compromise between both (all other tools). Overall, the choice of SNP caller strongly depends on which performance parameter should be maximized and whether ascertainment bias should be minimized to optimize downstream analysis, highlighting the need for studies that assess such differences.</p>
Illumina next generation ddRAD sequencing SNP data from: Contrasting genetic diversity and structure between endemic and widespread damselfishes are related to differing adaptive strategies
<p class="MsoNormal"><strong><u><span>Aim:</span></u></strong><span> Discerning when, where, and how processes of isolation lead to differing biogeography is especially complex for marine species with similar ecological niches and within the same geographic location. We assessed population genetics of congeneric and ecologically similar damselfishes within their overlapping distributions and across potential barriers to geneflow.</span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Taxon:</span></u></strong><span> <em>Dascyllus marginatus </em>(endemic) and <em>Dascyllus abudafur </em>(widespread)<em>.</em></span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Location:</span></u></strong><span> Coral reefs from the Red Sea, Djibouti, Yemen, Oman, and Madagascar. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Methods:</span></u></strong><span> We used RADseq derived SNPs to investigate key differences in population genetics between both species and discuss barriers shaping genetic differentiation (neutral vs. selective) and biogeography. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Results:</span></u></strong><strong><span> </span></strong><em><span>Dascyllus marginatus </span></em><span>inhabited the Red Sea, the coasts of Yemen (including Socotra), and the Gulf of Oman. <em>Dascyllus abudafur</em> species was present from the Red Sea to Madagascar but was absent from Yemen and Oman. Populations of <em>D. marginatus </em>had an order of magnitude higher genetic differentiation compared to <em>D. abudafur</em>, as well as several outlier loci (suggesting selective pressure), which were absent in <em>D. abudafur</em> despite equal sampling locations. In both species, specimens from the Red Sea and Djibouti formed one genetic cluster separated from all other locations. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Main conclusions:</span></u></strong><span> The stronger genetic structure at smaller geographic scale of the endemic species seems associated to faster adaptation to environmental differences; whereas the widespread species only experienced reduced geneflow and neutral differentiation at much larger geographic scales. Restrictive transitions (between the Gulf of Aqaba and the Red Sea or the Red Sea and the Gulf of Aden) did not affect the genetic architecture of either species, while the environmental shift within the Red Sea (at 22°N/20°N) affected the endemic but not the widespread species. Samples from continental Yemen revealed that a genetic break in the Gulf of Aden likely reflects historical colonization processes and not contemporary environmental regimes.</span></p>
Transitioning from microsatellites to SNP-based microhaplotypes in genetic monitoring programs: lessons from a 20-year time series of paired data.
<p>Many long-term genetic monitoring programs began before next-generation sequencing became widely available. Older programs can now transition to new marker systems usually consisting of 1000s of SNP loci, but there are still important questions about comparability, precision, and accuracy of key metrics estimated using SNPs. Ideally, transitioned programs should capitalize on new information without sacrificing continuity of inference across the time series. We combined existing microsatellite-based genetic monitoring information with SNP-based microhaplotypes obtained from archived samples of Rio Grande silvery minnow (<em>Hybognathus amarus</em>) across a 20-year time series to evaluate point estimates and trajectories of key genetic metrics. Demographic and genetic monitoring bracketed multiple collapses of the wild population, and included cases where captive-born repatriates comprised the majority of spawners in the wild. Even with smaller sample sizes, microhaplotypes yielded comparable and in some cases more precise estimates of variance genetic effective population size, multilocus heterozygosity and inbreeding compared to microsatellites because many more microhaplotype loci were available. Microhaplotypes also recorded shifts in allele frequencies associated with population bottlenecks. Trends in microhaplotype-based inbreeding metrics were associated with the fraction of hatchery-reared repatriates to the wild, and should be incorporated into future genomic monitoring. Although differences in accuracy and precision of some metrics were observed between marker types, biological inferences and management recommendations were consistent.</p>
Contrasting levels of hybridization across the two contact zones between two hedgehog species revealed by genome-wide SNP data
<p>Hybridization and introgression have played important roles in the history of various species, including lineage diversification and the evolution of adaptive traits. Hybridization can accelerate the development of reproductive isolation between diverging species, and thus valuable insight into the evolution of reproductive barrier formation may be gained by studying secondary contact zones. Hedgehogs of the genus <em>Erinaceus</em>, which are insectivores sensitive to changes in climate, are a pioneer model in Pleistocene phylogeography. The present study provides the first genome-wide SNP data regarding the <em>Erinaceus</em> hedgehogs species complex, offering a unique comparison of two secondary contact zones between <em>Erinaceus</em> <em>europaeus</em> and <em>E</em>. <em>roumanicus</em>. Results confirmed diversification of the genus during the Pleistocene period and detected a new refugial lineage of <em>E</em>. <em>roumanicus</em> outside the Mediterranean region, most likely in the Ponto-Caspian region. In the Central European zone, the level of hybridization was low, whereas in the Russian-Baltic zone, both species hybridise extensively. Asymmetrical gene flow from <em>E</em>. <em>europaeus</em> to <em>E</em>. <em>roumanicus</em> suggests that reproductive isolation varies according to the direction of the crosses in the hybrid zones. However, no loci with significantly different patterns of introgression were detected. Markedly different pre- and post-zygotic barriers, and thus diverse modes of species boundary maintenance in the two contact zones, likely exist. This pattern is probably a consequence of the different ages and thus of the different stages of evolution of reproductive isolating mechanisms in each hybrid zone.</p>
Speciation in coastal basins driven by staggered headwater captures: Dispersal of a species complex, Leporinus bahiensis, as revealed by genome-wide SNP data
<p>Past sea level changes and geological instability along watershed boundaries have largely influenced fish distribution across coastal basins, either by dispersal via palaeodrainages now submerged or by headwater captures, respectively. Accordingly, the South American Atlantic coast encompasses several small and isolated drainages that share a similar species composition, representing a suitable model to infer historical processes. <em>Leporinus</em> <em>bahiensis</em> is a freshwater fish species widespread along adjacent coastal basins over narrow continental shelf with no evidence of palaeodrainage connections at low sea level periods. Therefore, this study aimed to reconstruct its evolutionary history to infer the role of headwater captures in the dispersal process. To accomplish this, we employed molecular-level phylogenetic and population structure analyses based on Sanger sequences (5 genes) and genome-wide SNP data. Phylogenetic trees based on Sanger data were inconclusive, but SNPs data did support the monophyletic status of <em>L. bahiensis</em>. Both COI and SNP data revealed structured populations according to each hydrographic basin. Species delimitation analyses revealed from 3 (COI) to 5 (multilocus approach) MOTUs, corresponding to the sampled basins. An intricate biogeographic scenario was inferred and supported by Approximate Bayesian Computation (ABC) analysis. Specifically, a staggered pattern was revealed and characterized by sequential headwater captures from basins adjacent to upland drainages into small coastal basins at different periods. These headwater captures resulted in dispersal throughout contiguous coastal basins, followed by deep genetic divergence among lineages. To decipher such recent divergences, as herein represented by <em>L. bahiensis </em>populations, we used genome-wide SNPs data. Indeed, the combined use of genome-wide SNPs data and ABC method allowed us to reconstruct the evolutionary history and speciation of <em>L. bahiensis</em>. This framework might be useful in disentangling the diversification process in other neotropical fishes subject to a reticulate geological history. </p>
Genetic diversity and population structure from a Peruvian nucleus cattle herd using SNP data
<p>New-generation sequencing technologies, among them SNP chips for massive genotyping, have proven to be useful for the effective management of genetic resources. Also, developing nucleus herds is an effective method for genetic improvement work. To date, molecular studies in Peruvian cattle are still in their infancy. To close this gap, we here employed two SNP panels (BovineHD and Bovine100K) to determine for the first time the Peruvian nucleus herd's genetic diversity and population structure that belong to INIA. This nucleus comprises Brahman (N=16), Braunvieh (N=14), Gyr (N=11), and Fleckvieh (N=22) breeds. Additionally, samples from a locally adapted creole cattle, the Arequipa Fighting Bull (AFB, N=12), were incorporated into the study. The genetic diversity indices in all breeds showed a high proportion of polymorphic SNPs, varying from 69.37% in Gyr to 80.81% in Braunvieh. Also, Braunvieh possessed the highest observed heterozygosity (0.53±0.17), while Brahman possessed the lowest (0.44±0.10), indicating that the former is more diverse compared to the other cattle breed groups. According to the molecular variance analysis, 83.92% of the variance occurs within individuals, whereas 16.0% occurs between populations. The pairwise FST estimates between breeds showed values that ranged from 0.054 (Braunvieh vs AFB) to 0.266 (Brahman vs AFB). Pairwise Reynold's distance showed a pattern similar to the one obtained with the FST statistics, with values ranging from 0.058 to 0.309. A dendrogram was constructed using the Neighbor-Joining clustering algorithm, and similar to the principal coordinate analysis, three groups were identified. Results showed a clear separation between <em>Bos</em> <em>indicus</em> (Brahman and Gyr) and <em>B</em>. <em>taurus</em> breeds (Braunvieh and Fleckvieh). For Fleckvieh and Braunvieh, there were two subgroups each one of them grouping with the AFB group. Similar results were obtained with ADMIXTURE analysis with K= 3 as the most optimal number for the inferred genetic structure of the populations. The results from the current study would contribute to the appropriate management avoiding loss of genetic variability in these breeds and to future improvements for this nucleus. Additional work is needed to speed up the breeding process in the Peruvian cattle system.</p>
Obuasi case study data: Performance of neutral SNP barcodes to determine genetic diversity and structure of Plasmodium falciparum in Africa
<p>A small number of informative biallelic single nucleotide polymorphisms (SNPs) have been proposed to be an economical method to fast-track the genotyping and relatedness analysis of <em>Plasmodium</em> <em>falciparum</em> in malaria-endemic areas. Whilst used successfully in low-transmission areas where infections are monoclonal and highly related, we present the first study to evaluate the performance of these 24- and 96-SNP molecular barcodes in African countries characterised by moderate-to-high transmission. Using haplotypes generated from the MalariaGEN <em>P. falciparum</em> Community Project version 6 database, 52.3% of infections were multiclonal, generating high frequencies of mixed-allele calls (MACs) per isolate. Both multiclonality and low heterozygosity of SNPs impeded haplotype construction for analyses of relatedness. Although fewer SNPs provided usable data, these SNP barcodes weakly identified genetic differentiation across large geographic distances. However, both minor and major alleles' frequencies were temporally unstable. We conclude that these standardised SNP barcodes are vulnerable to ascertainment bias. While large numbers of SNPs acquired by whole-genome sequencing and computational methods to construct haplotypes present a way forward, these approaches may not be practical or cost-effective for surveillance on large scales in malaria-endemic areas. </p>
Sample extraction and SNP sequencing data for: Identification of sex-linked SNP markers in wild populations of monomorphic birds
<p><span>Single-nucleotide polymorphism (SNP) analyses are a powerful tool for population genetics, pedigree reconstruction and phenotypic trait mapping. However, the untapped potential of SNP markers to discriminate the sex of individuals in species with reduced sexual dimorphism or of individuals during immature stages remains a largely unexplored avenue. Here, we develop a novel protocol for molecular sexing of birds based on the detection of unique Z- and W-linked SNP markers. Our method is based on the identification of two unique loci, one in each sexual chromosome. Individuals are considered males when they show no calls for the W-linked SNP and are heterozygotic or homozygotic for the Z-linked SNP, while females show both Z- and W-linked SNP calls. We validated the method in the Jackdaw (<em>Corvus</em> <em>monedula</em>). The reduced sexual dimorphism in this species makes it difficult to sex individuals in the wild. We assessed the reliability of the method using 36 individuals of known sex and found that their sex was correctly assigned in 100% of cases. The sex-linked markers also proved to be widely applicable to discriminate males and females from a sample of 927 genotyped individuals of different maturity stages with an accuracy of 99.5%. Given that SNP markers are increasingly used in quantitative genetic analyses of wild populations, the approach we propose has a great potential to be integrated into broader genetic research programmes without the need for additional sexing techniques.</span></p>
Pakistani historical wheat panel 37K SNP data
<p>A collection of 196 historical wheat cultivars of Pakistan released between 1911 to 2022 were subjected to DNA fingerprinting using 16K genotyping-by-targeted sequencing (GBTS) platform. This platform is based on NGS and the 16K probes were resequenced. The resequencing data was aligned to Chinese Spring RefSeq version 1.1 and SNPcalling was performed. This resulted in 37K mSNP (multiple SNPs) DNA fingerprinting data. This is thus far the most comprehensive DNA fingerprinting data of all released cultivars of wheat so far. This data is publically available and can be used in research and publications subject to the acknowledgement. </p>
Data from: Recommendations for population and individual diagnostic SNP selection in non-model species
Open the record for dataset details and reuse information.
Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs
Open the record for dataset details and reuse information.
Speciation in coastal basins driven by staggered headwater captures: Dispersal of a species complex, Leporinus bahiensis, as revealed by genome-wide SNP data
Open the record for dataset details and reuse information.
Data from high throughput SNP-chip as cost effective new monitoring tool for assessing invasion dynamics in the comb jelly Mnemiopsis leidyi
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.