Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
320
datasets available to search
ShareScore release 0.9.0
Dataset results
320 results for “SNP array”
Data for "How Array Design Creates SNP Ascertainement Bias"
<p>The repository contains the raw SNP data in vcf format for the publication "How Array Design Creates SNP Ascertainment Bias". Note that the variants are <strong>not</strong> filtered at this timepoint. Samples starting with pl_ are pooled sequences of ~10 individuals while samples starting with i_ were individually sequenced. Please find detailed information about samples, raw sequencing data and SNP calling pipeline in the linked preprint (<a href="https://doi.org/10.1101/833541">https://doi.org/10.1101/833541</a>)/ publication (<a href="https://doi.org/10.1371/journal.pone.0245178">https://doi.org/10.1371/journal.pone.0245178</a>). In case you need additional information, please contact<a href="mailto:johannes.geibel@uni-goettingen.de"> johannes.geibel@uni-goettingen.de</a></p>
Genotyping of the Chinese Spring x Renan mapping population with the TaBW280K SNP array
<p>The TaBW280K SNP array (Rimbert et al., PLoS ONE 2018) was used to genotype 430 Single Seed Descent (SSD) individuals<br> derived from a cross between Chinese Spring and Renan (CsRe; Choulet et al., Science 2014). Out of the 280,226 probesets, 85,276 were found to be polymorphic between the two parental lines and PHR on the population. Eventually, 83,721 (98.2%) SNPs were genetically mapped in 21 linkage groups corresponding to the 21 chromosomes of bread wheat, with no unlinked markers. This file contains the genotyping data of the 430 SSD lines.<br> </p>
Rye600K SNP Array 'Lo7' Map
<p>Mapping position of rye600K SNP array markers developed by Bauer., et al. 2017, on the reference chromosomal-scale rye reference genome 'Lo7'. Shared in the hope that it provides a valuable and easily-accessible resource for further studies in rye genomics.</p> <p><strong>Method:</strong></p> <p>Positional data of 600K SNP markers was obtained by mapping each of the 600843 SNP marker sequences to the rye reference genome ‘Lo7’ using NCBI blastn (v. 2.9.0+) function. Mapping position of SNPs were hereafter stringently filtered for I) complete SNP sequence alignment, and II) maximum of 1 mismatch to ensure an accurate positioning. </p> <p> </p> <p> </p>
Development of a high-density 665 K SNP array for rainbow trout genome-wide genotyping. Supplemental VCF file
<p>Single nucleotide polymorphism (SNP) arrays, also named « SNP chips », enable very large numbers of individuals to be genotyped at a targeted set of thousands of genome-wide identified markers. We used preexisting variant datasets from USDA, a French commercial line and 30X-coverage whole genome sequencing of INRAE isogenic lines to develop an Affymetrix 665 K SNP array (HD chip) for rainbow trout. In total, we identified 32,372,492 SNPs that were polymorphic in the USDA or INRAE databases. A subset of identified SNPs were selected for inclusion on the chip, prioritizing SNPs whose flanking sequence uniquely aligned to the Swanson reference genome, with homogenous repartition over the genome and the highest Minimum Allele Frequency in both USDA and French databases. Of the 664,531 SNPs which passed the Affymetrix quality filters and were manufactured on the HD chip, 65.3% and 60.9% passed filtering metrics and were polymorphic in two other distinct French commercial populations in which, respectively, 288 and 175 sampled fish were genotyped. Only 576,118 SNPs mapped uniquely on both Swanson and Arlee reference genomes, and 12,071 SNPs did not map at all on the Arlee reference genome. Among those 576,118 SNPs, 38,948 SNPs were kept from the commercially available medium-density 57K SNP chip. We demonstrate the utility of the HD chip by describing the high rates of linkage disequilibrium at 2 kb to 10 kb in the rainbow trout genome in comparison to the linkage disequilibrium observed at 50 kb to 100 kb which are usual distances between markers of the medium-density chip.</p> <p> </p> <p>File submitted correspond to the supplementary data 1 of the publication (under submission) : INRAE_USDA_MAF1.vcf.gz</p>
70K SNP array data for Lumpfish (Cyclopterus lumpus) across the trans-Atlantic
<p>In marine species with large populations and high dispersal potential, large-scale genetic differences and clinal trends in allele frequency can provide insight into the evolutionary processes that shape diversity. Lumpfish, <em>Cyclopterus lumpus</em>, is found throughout the North Atlantic and has traditionally been harvested for roe and more recently used as a cleaner fish in salmon aquaculture. We used a 70K SNP array to evaluate trans-Atlantic differentiation, genetic structuring, and clinal variation across the North Atlantic. Basin-scale structuring between the Northeast and Northwest Atlantic was significant, with enrichment for loci associated with developmental/mitochondrial function. We identified a putative structural variant on chromosome 2, likely contributing to differentiation between Northeast and Northwest Atlantic Lumpfish, and consistent with post-glacial trans-Atlantic secondary contact. Redundancy Analysis identified climate associations both in the Northeast (<em>N</em> = 1269 loci) and Northwest (<em>N</em> = 1637 loci), with 103 shared loci between them. Clinal patterns in allele frequencies were observed in some loci (15% - Northwest and 5% - Northeast) of which 708 loci were shared and involved with growth, developmental processes, and locomotion. The combined evidence of trans-Atlantic differentiation, environmental associations, and clinal loci, suggests that both regional and large-scale potentially-adaptive population structuring is present across the North Atlantic.</p>
70K SNP array data for Lumpfish (Cyclopterus lumpus) across the trans-Atlantic
Open the record for dataset details and reuse information.
Aedes aegypti in North America (Microsatellite and SNP array)
<p>The <em>Aedes aegypti</em> mosquito first invaded the Americas about 500 years ago and today is a widely distributed invasive species and the primary vector for viruses causing dengue, chikungunya, Zika, and yellow fever. Here we test the hypothesis that the North American colonization by <em>Ae. aegypti</em> occurred via a series of founder events. We present findings on genetic diversity, structure, and demographic history using data from 70 <em>Ae. aegypti</em> populations in North America genotyped at 12 microsatellite loci and/or ~20,000 single nucleotide polymorphisms (SNPs), the largest genetic study of the region to date. We find evidence consistent with a colonization driven by serial founder effect (SFE), with Florida as the putative source for a series of westward invasions. This scenario was supported by 1) a decrease in the genetic diversity of <em>Ae. aegypti </em>populations moving west, 2) a correlation between pairwise genetic and geographic distances, and 3) demographic analysis based on allele frequencies. A few <em>Ae. aegypti</em> populations on the west coast do not follow the general trend, likely due to a recent and distinct invasion history. We argue that SFE provides a helpful albeit simplified model for the movement of <em>Ae. aegypti </em>across North America, with outlier populations warranting further investigation.</p>
Determining haploblocks and haplotypes in the MAGIC winter wheat population WM-800 based on the wheat 15k Infinium and the 135k Affymetrix SNP arrays
<p><span>Haplotypes are derived from single nucleotide polymorphisms (SNPs). They are beneficial (i) to remove redundant sequence information in genetic populations and, more important, (ii) to distinguish more than two variants/alleles at a genomic locus. A haploblock locus, made of multiple haplotypes, is very useful in multiparent-advanced-generation-intercross (MAGIC) populations, where, ideally, multiple founder alleles need to be distinguished at each locus to subsequently carry out efficient genome-wide association analysis studies (GWAS). </span></p> <p><span>In this regard, the dataset contains genotype matrices (made of SNP, haploblock and haplotype data) for 800 lines of the MAGIC WHEAT population WM-800 (Sannemann et al. 2018). The datasets are based on genotyping the lines with both the already published wheat 15k Infinium SNP array (Sannemann et al. 2018) and the new wheat 135k Affymetrix SNP array.</span></p>
(SNP Array) Single-Cell Multi-Omics Identifies Chronic Inflammation as a Driver of TP53 mutant Leukaemic Evolution
<p>Single nucleotide polymorphism (SNP) array data files related our publication titled "Single-Cell Multi-Omics Identifies Chronic Inflammation as a Driver of <em>TP53 </em>mutant Leukaemic Evolution".</p>
Aedes aegypti in North America (Microsatellite and SNP array)
Open the record for dataset details and reuse information.
Determining haploblocks and haplotypes in the MAGIC winter wheat population WM-800 based on the wheat 15k Infinium and the 135k Affymetrix SNP arrays
Open the record for dataset details and reuse information.
SNP array for parentage assignment of the Manila clam, Ruditapes philippinarum
<p>The Manila clam <i>Ruditapes philippinarum</i>, a major cultured shellfish species, is threatened by infection with the microparasite <i>Perkinsus olseni</i>, whose prevalence increases with high water temperatures. Under the current trend of climate change, the already severe effects of this parasitic infection might rapidly increase the frequency of mass mortality events. Treating infectious diseases in bivalves is notoriously problematic, therefore selective breeding for resistance represents a key strategy for mitigating the negative impact of pathogens. A crucial step in initiating selective breeding is the estimation of genetic parameters for traits of interest, which relies on the ability to record parentage and accurate phenotypes in a large number of individuals. Here, to estimate the heritability of resistance against <i>P. olseni</i>, a field experiment mirroring conditions in industrial clam production was set up, a genomic tool was developed for parentage assignment, and parasite load was determined through quantitative PCR.</p> <p>A mixed-family cohort of potentially 1479 clam families was produced in a hatchery by mass spawning of 53 dams and 57 sires. The progenies were seeded in a commercial clam production area in the Venice lagoon, Italy, where high prevalence of <i>P. olseni</i> had previously been reported. Growth and parasite load were monitored every month and, after one year, more than 1000 individuals were collected and DNA and phenotype records.</p> <p>A 245-SNP panel was developed using candidate markers obtained from a pooled sequencing approach on two DNA samples from all the potential parents and from a Venice lagoon clam population. For 246 individuals of the mixed-family F1, sire and dam representation were high (75 and 85%, respectively), indicating a very limited risk of inbreeding. Moderate heritability (0.20 – 0.30) was estimated for growth traits, while parasite load showed high heritability, estimated at 0.52. No significant genetic correlations were found between growth-associated traits and parasite load.</p> <p>Overall, the study shows high potential for selecting clams resistant to parasite<i> </i>load<i>.</i> Breeding for resistance may help limit the negative effects of climate change on clam production, as the prevalence of the parasite is predicted to increase under a future scenario of higher temperatures. Finally, the limited genetic correlation between resistance and growth suggests that breeding programs could incorporate dual selection without negative interactions.</p>
Data from: Transatlantic secondary contact in Atlantic salmon, comparing microsatellites, a SNP array, and Restriction Associated DNA sequencing for the resolution of complex spatial structure
Identification of discrete and unique assemblages of individuals or populations is central to the management of exploited species. Advances in population genomics provide new opportunities for re-evaluating existing conservation units but comparisons among approaches remain rare. We compare the utility of RAD-seq, a single nucleotide polymorphism (SNP) array and a microsatellite panel to resolve spatial structuring under a scenario of possible trans-Atlantic secondary contact in a threatened Atlantic Salmon, Salmo salar, population in southern Newfoundland. Bayesian clustering indentified two large groups subdividing the existing conservation unit and multivariate analyses indicated significant similarity in spatial structuring among the three data sets. mtDNA alleles diagnostic for European ancestry displayed increased frequency in southeastern Newfoundland and were correlated with spatial structure in all marker types. Evidence consistent with introgression among these two groups was present in both SNP data sets but not the microsatellite data. Asymmetry in the degree of introgression was also apparent in SNP data sets with evidence of gene flow towards the east or European type. This work highlights the utility of RAD-seq based approaches for the resolution of complex spatial patterns, resolves a region of trans-Atlantic secondary contact in Atlantic Salmon in Newfoundland and demonstrates the utility of multiple marker comparisons in identifying dynamics of introgression.
Data from: SNP-array reveals genome wide patterns of geographical and potential adaptive divergence across the natural range of Atlantic salmon (Salmo salar)
Atlantic salmon (Salmo salar) is one of the most extensively studied fish species in the world due to its significance in aquaculture, fisheries and ongoing conservation efforts to protect declining populations. Yet, limited genomic resources have hampered our understanding of genetic architecture in the species and the genetic basis of adaptation to the wide range of natural and artificial environments it occupies. In this paper, we describe the development of a medium density Atlantic salmon SNP-array based on Expressed Sequence Tags (ESTs) and genomic sequencing. The array was used in the most extensive assessment of population genetic structure performed to date in this species. A total of 6176 informative SNPs were successfully genotyped in 38 anadromous and freshwater wild populations distributed across the species natural range. Principal component analysis clearly differentiated European and North American populations, and within Europe, three major regional genetic groups were identified for the first time in a single analysis. We assessed the potential for the array to disentangle neutral and putative adaptive divergence of SNP allele frequencies across populations and among regional groups. In Europe, secondary contact zones were identified between major clusters where endogenous and exogenous barriers could be associated, rendering the interpretation of environmental influence on potentially adaptive divergence equivocal. A small number of markers highly divergent in allele frequencies (outliers) were observed between (multiple) freshwater and anadromous populations, between northern and southern latitudes, and when comparing Baltic populations to all others. We also discuss the potential future applications of the SNP-array for conservation, management and aquaculture.
Data from: Discovery of 20,000 RAD–SNPs and development of a 52-SNP array for monitoring river otters
Many North American river otter (Lontra canadensis) populations are threatened or recovering but are difficult to study because they occur at low densities, it is difficult to visually identify individuals, and they inhabit aquatic environments that accelerate degradation of biological samples. Single nucleotide polymorphisms (SNPs) can improve our ability to monitor demographic and genetic parameters of difficult to study species. We used restriction site associated DNA (RAD) sequencing to discover 20,772 SNPs present in Montana, USA, river otter populations, including 14,512 loci that were also variable in at least one other population range-wide. After applying careful filtering criteria meant to minimize ascertainment bias and identify high quality, highly heterozygous (H o = 0.2–0.50) SNPs, we developed and tested 52 independent SNP qPCR genotyping assays, including 41 that performed well with diluted DNA. The 41 loci provided high power for population assignment tests with only 1 misassignment (1.6 %) between closely neighboring populations. Our SNPs showed high power to differentiate individuals and assign them to population of origin, as well as strong concordance of genotypes from high and diluted concentrations of DNA, and between original RAD and the SNP qPCR array.
Data from: Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks
An Illumina Infinium SNP genotyping array was constructed for European white oaks. Six individuals of Quercus petraea and Q. robur were considered for SNP discovery using both previously obtained Sanger sequences across 676 gene regions (1371 in vitro SNPs) and Roche 454 technology sequences from 5112 contigs (6542 putative in silico SNPs). The 7913 SNPs were genotyped across the six parental individuals, full-sib progenies (one within each species and two interspecific crosses between Q. petraea and Q. robur) and three natural populations from south-western France that included two additional interfertile white oak species (Q. pubescens and Q. pyrenaica). The genotyping success rate in mapping populations was 80.4% overall and 72.4% for polymorphic SNPs. In natural populations, these figures were lower (54.8% and 51.9%, respectively). Illumina genotype clusters with compression (shift of clusters on the normalized x-axis) were detected in ~25% of the successfully genotyped SNPs and may be due to the presence of paralogues. Compressed clusters were significantly more frequent for SNPs showing a priori incorrect Illumina genotypes, suggesting that they should be considered with caution or discarded. Altogether, these results show a high experimental error rate for the Infinium array (between 15% and 20% of SNPs potentially unreliable and 10% when excluding all compressed clusters), and recommendations are proposed when applying this type of high-throughput technique. Finally, results on diversity levels and shared polymorphisms across targeted white oaks and more distant species of the Quercus genus are discussed, and perspectives for future comparative studies are proposed.
Data from: Transatlantic secondary contact in Atlantic salmon, comparing microsatellites, a SNP array, and Restriction Associated DNA sequencing for the resolution of complex spatial structure
Open the record for dataset details and reuse information.
Data from: Discovery of 20,000 RAD–SNPs and development of a 52-SNP array for monitoring river otters
Open the record for dataset details and reuse information.
Data from: New resources for genetic studies in Populus nigra: genome wide SNP discovery and development of a 12k Infinium array
Open the record for dataset details and reuse information.
Data from: Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.