Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
123
datasets available to search
ShareScore release 0.9.0
Dataset results
123 results for “SNP genotyping”
Construction of genetic linkage map based on SNP markers, QTL mapping and detection of candidate genes of growth-related traits in Pacific abalone using genotyping-by-sequencing
<p><a name="_Hlk72585736"><span>Pacific abalone (<i>Haliotis discus hannai</i>) is a commercially important high valued molluscan species. Its wild population has decreased in recent years. Pacific abalone is widely cultured in Korea. Traditional breeding programs have been implemented for hatchery production of abalone seeds. To obtain more genetic information for the molecular breeding program, a high-density linkage map and quantitative trait locus (QTL) for three growth-related traits was constructed for Pacific abalone. F1 cross population with two parents were sampled to construct the linkage map using genotyping by sequencing (GBS). A total of 664,630,534 clean reads and 56,686 SNPs were generated. In sum, 3,345 segregating SNPs were used to construct a consensus linkage map. The map spanned 1,747.023 cM with 18 linkage groups and an average interval of 0.55 cM. QTL analysis revealed two significant QTL in LG10 on the consensus linkage map in each growth-related trait. Both the QTLs are located in the telomere region of the chromosome. Moreover, four potential candidate genes for growth-related traits were identified in the QTL region. Expression analysis revealed that identified genes are involved in growth regulation of abalone. The newly constructed genetic linkage map, growth-related QTLs and potential candidate genes identified in the present study can be used as valuable genetic resources and will be useful for marker-assisted selection (MAS) of Pacific abalone in molecular breeding program.</span></a></p>
Data from: A high density SNP chip for genotyping great tit (Parus major) populations and its application to studying the genetic architecture of exploration behaviour
High density SNP microarrays ('SNP chips') are a rapid, accurate and efficient method for genotyping several hundred thousand polymorphisms in large numbers of individuals. While SNP chips are routinely used in human genetics and in animal and plant breeding, they are less widely used in evolutionary and ecological research. In this paper we describe the development and application of a high density Affymetrix Axiom chip with around 500 000 SNPs, designed to perform genomics studies of great tit (Parus major) populations. We demonstrate that the per-SNP genotype error rate is well below 1% and that the chip can also be used to identify structural or copy number variation (CNVs). The chip is used to explore the genetic architecture of exploration behaviour (EB), a personality trait that has been widely studied in great tits and other species. No SNPs reached genome-wide significance, including at DRD4, a candidate gene. However, EB is heritable and appears to have a polygenic architecture. Researchers developing similar SNP chips may note: (i) SNPs previously typed on alternative platforms are more likely to be converted to working assays, (ii) detecting SNPs by more than one pipeline, and in independent datasets, ensures a high proportion of working assays, (iii) allele frequency ascertainment bias is minimised by performing SNP discovery in individuals from multiple populations and (iv) samples with the lowest call rates tend to also have the greatest genotyping error rates.
Data from: Multiplex preamplification PCR and microsatellite validation allows accurate single nucleotide polymorphism (SNP) genotyping of historical fish scales
Incorporating historical tissues into the study of ecological, conservation, and management questions can broaden the scope of population genetic research by enhancing our understanding of evolutionary processes and anthropogenic influences on natural populations. Genotyping historical and low-quality samples has been plagued by challenges associated with low amounts of template DNA and the potential for preexisting DNA contamination among samples. We describe a two-step process designed to (i) accurately genotype large numbers of historical low-quality scale samples in a high-throughput format and (ii) screen samples for preexisting DNA contamination. First, we describe how an efficient multiplex preamplification PCR of 45 single nucleotide polymorphisms (SNPs) can generate highly accurate genotypes with low failure and error rates in subsequent SNP genotyping reactions of individual historical scales from sockeye salmon (Oncorhynchus nerka). Second, we demonstrate how the method can be modified for the amplification of microsatellite loci to detect preexisting DNA contamination. A total of 760 individual historical scale and 182 contemporary fin clip samples were genotyped and screened for contamination. Genotyping failure and error rates were exceedingly low and similar for both historical and contemporary samples. Preexisting contamination in 21% of the historical samples was successfully identified by screening the amplified microsatellite loci. The potential for automation, low failure and error rates, and ability to multiplex both the preamplification and subsequent genotyping reactions combine to make the protocol ideally suited for efficiently genotyping large numbers of potentially contaminated low-quality sources of DNA.
SNP genotypes for healthy and CIM-affected GSDs
<p>German shepherd dogs (GSDs) are predisposed to an inherited motility disorder of the esophagus, termed congenital idiopathic megaesophagus (CIM), in which swallowing is ineffective and the esophagus is enlarged. Affected puppies are unable to properly pass food into their stomachs and consequently regurgitate their meals and show a failure to thrive, often leading to euthanasia. Here, we generated genome-wide SNP profiles for healthy and CIM-affected GSDs using the Illumina CanineHD BeadChip, containing 220,853 SNPs.</p>
Impatiens glandulifera SNP and SilicoDArT genotyping data
<p>We conducted genomic characterization based on SNP and SilicoDArT markers on the invasive Himalayan balsam (<i>Impatiens glandulifera</i>) plants originating from the native and non-native regions of their distribution. When genetic relationships were explored by PCoA based on SNP and SilicoDArT marker data, the first, second and third principal coordinates explained altogether 37.4% and 31.0% of the variability, respectively. Samples from the UK, Canada and Pakistan grouped together, while Indian plants were clearly distinct based on SNP markers but relatively close to the UK-Canada-Pakistan group based on SilicoDArT markers. Constructed trees differentiated the individuals into clusters resembling the patterns observed by PCoA.<span> The Bayesian BAPS analysis revealed that the individuals were distributed in seven clusters, representing samples from each of the four Finnish populations, India, Pakistan and the combination of the UK and Canada. Similar clustering was visible in the constructed UPGMA tree. The Indian cluster did not display any ancestral gene flow with the others, while the Pakistani cluster showed ancestral gene flow only with the combined UK and Canada cluster. Furthermore, the latter cluster displayed ancestral gene flow with the Finnish populations varying from 0% to 3.1%. The AMOVA analysis showed that 45% and 26% of genetic variation was present among the <i>I. glandulifera</i> groups/populations and the rest within them based on SNP and SilicoDArT markers, respectively. Overall, the Bayesian BAPS analysis</span> <span>and the following gene flow network were the most informative tools for resolving relationships among native and introduced plants. </span></p>
Western redcedar single nucleotide polymorphism (SNP) genotyping data for genomic selection and population genetics
<p>Western redcedar (<em>Thuja plicata</em>) Single Nucleotide Polymorphism (SNP) data in Variant Call Format (VCF) for genomic selection training and target populations, genomic selection parents, and self-fertilized (selfing) lines, comprising 4,833 trees.</p> <p>Targeted sequencing-based genotyping was done by Capture-Seq methodology at Rapid Genomics (Neves est al. 2013). A set of 57,000 probes as designed for initial marker discovery, from which a panel of 20,858 probes was selected for genotyping. A set of transcriptomes (Shalev et al. 2018) (PRJNA704616) was aligned to the reference genome to identify SNPs. Candidate probes (120 nt) were initially designed in silico and 57,000 selected by removing candidates with poor base composition for hybridization (GC content <0.2 and >0.6, high G content >0.2 and long homopolymers >7), followed by removing probes aligning to more than one position on the reference genome (≥90% identity and length). The 57,000 probes represent 14,517 scaffolds (average 3.9 probes/scaffold), with 37,275 targeting at least one SNP and 19,725 mapping to intergenic regions not containing pre-identified SNPs. A set of 128 individuals were selected to validate the 57,000 probe panel and associated polymorphisms. Genomic DNA (0.5 ug) was fragmented (mean size 300 bp), followed by repair of ends, phosphorylation, adenylation, ligation of Illumina compatible adapters containing 8bp indexes and 5’ T-overhang, and 10 cycles PCR amplification with universal primers to produce sequencing-ready libraries. Libraries were quantified using PicoGreen. Libraries from 16 samples were pooled, hybridized to the 120 nt RNA probes following Agilent’s SureSelect Target Enrichment System (Agilent Technologies) and sequenced on an Illumina HiSeq X machine with paired-end 150bp cycle for an average sequencing depth per sample of 15X. Sequence data were aligned to the reference genome with BWA-MEM (http://arxiv.org/abs/1303.3997) and sets of four samples were combined to increase sequencing depth for identifying markers. Putative SNPs were identified using Freebayes (http://arxiv.org/abs/1207.3907) in 150bp on either side of the 57,000 probes and filtered probes that had more than 17 SNPs per 420 bp target region (150bp + 120bp + 150bp). The sequencing depth of the probes was used to select the final set of 20,885 probes, removing probes on both sides of the distribution (low and high sequencing depth), for Capture-Seq on the remainder of the samples.</p>
Spatiotemporal monitoring of the rare Northern dragonhead, Dracocephalum ruyschiana (Lamiaceae): SNP genotyping and environmental niche modelling herbarium specimens
<p><strong>Aim: </strong>We have studied spatiotemporal genetic change in the Northern dragonhead, a plant species that has experienced a drastic population decline and habitat loss in Europe. We add a temporal perspective to the monitoring of dragonhead in Norway by genotyping herbarium specimens up to 200 years old. We also assess whether dragonhead has achieved its potential distribution in Norway. Location: Europe (mainly Norway)</p> <p><strong>Methods:</strong> We have applied a microfluidic array consisting of 96 SNP markers on 130 herbarium specimens collected from 1820 to 2008, mainly from Norway (83) but also beyond (47). We have compared our new genotype data with existing data from modern samples. We have modelled the species' environmental niche and potential distribution in Norway using sample metadata and observational records.</p> <p><strong>Results: </strong>The SNP array successfully genotyped all included herbarium specimens. The captured genetic diversity was negatively correlated with distance from Norway. The historical-modern comparison revealed similar genetic structure and diversity across space and limited genetic change through time in Norway. The ENM suggests that dragonhead is anchored in warmer and drier habitats.</p> <p><strong>Main conclusions: </strong>With appropriate design procedures, the SNP array technology is promising for genotyping old herbarium specimens. We found no signs of any regional bottleneck. The regional areas in Norway have remained genetically divergent, however, both from each other and more so from populations outside of Norway, rendering continued protection of the species in Norway relevant. The ENM suggests that dragonhead has not fully achieved its potential distribution in Norway.</p>
SSR and SNP profiles obtained for 40 non-redundant grapevine genotypes found in the living collection of Ain Taoujdate (Morocco)
<p>This dataset includes the genetic profiles (13 SSR and 240 SNP markers) of 40 grapevine genotypes identified in the living collection of Ain Taoujdate (Morocco)</p>
SNP genotyping of Lord Howe woodhen (Hypotaenidia sylvestris) from museum skins and contemporary blood samples
<p>These data are from a conservation genetics project investigating population structure, dispersal and genetic bottlenecks in the Lord Howe woodhen <i>Hypotaenidia sylvestris.</i> This species recovered from near extinction in the 1970s to approximately 250 individuals in 2017. We used single nucleotide polymorphisms (SNPs) to genotype samples of both the contemporary population and 100-year-old museum specimens. We discovered strong population structuring between mountain and lowland "subpopulations" in both the contemporary and historic populations. This is indicative of restricted dispersal at fine spatial scales associated with rugged topography. There was also a decline in genetic diversity over the past century. We recommend ongoing genetic monitoring and translocations to increase genetic diversity within the re-established lowland subpopulation which although numerically stronger is genetically depleted relative to the mountain subpopulation.</p>
Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements
Sample availability limits population genetics research on many species, especially taxa from regions with high diversity. However, many such species are well represented in museum collections assembled before the molecular era. Development of techniques to recover genetic data from these invaluable specimens will benefit biodiversity science. Using a mixture of freshly preserved and historical tissue samples, and a sequence capture probe set targeting >5000 loci, we produced high-confidence genotype calls on thousands of single nucleotide polymorphisms (SNPs) in each of five South-East Asian bird species and their close relatives (N = 27–43). On average, 66.2% of the reads mapped to the pseudo-reference genome of each species. Of these mapped reads, an average of 52.7% was identified as PCR or optical duplicates. We achieved deeper effective sequencing for historical samples (122.7×) compared to modern samples (23.5×). The number of nucleotide sites with at least 8× sequencing depth was high, with averages ranging from 0.89 × 106 bp (Arachnothera, modern samples) to 1.98 × 106 bp (Stachyris, modern samples). Linear regression revealed that the amount of sequence data obtained from each historical sample (represented by per cent of the pseudo-reference genome recovered with ≥8× sequencing depth) was positively and significantly (P ≤ 0.013) related to how recently the sample was collected. We observed characteristic post-mortem damage in the DNA of historical samples. However, we were able to reduce the error rate significantly by truncating ends of reads during read mapping (local alignment) and conducting stringent SNP and genotype filtering.
SNP genotyping of North Head and northern Sydney Long-nosed bandicoots (Perameles nasuta)
<p>Wildlife species impacted by habitat loss and fragmentation often require conservation efforts to maintain populations. Long-nosed bandicoots (<i>Perameles nasuta</i>) still persist within the highly urbanised matrix of northern Sydney (Australia). These data are from a conservation genetics project investigating population structure and genetic diversity of the North Head Long-nosed bandicoot (<em>Perameles nasuta) </em>population and individuals from surrounding suburbs throughout northern Sydney.</p> <p>The population at North Head, Sydney, is currently listed as an <i>Endangered</i> population due to its small size, apparent isolation and other threats. To support future management, we used 1,446 single nucleotide polymorphism markers (SNPs) from 167 bandicoots to: i) assess the assumption of isolation and determine if genetic structuring is present between North Head and individuals from 11 other localities in northern Sydney, and ii) investigate genetic diversity over time in the North Head population from 2002 to 2018. Analyses confirmed population structuring and genetic divergence between North Head and greater northern Sydney. Three distinct populations were identified that corresponded to geographic localities (North Head, northern Sydney and Mosman). All populations were significantly differentiated (<i>F</i><sub>ST</sub> = 0.171–0.345), suggesting local genetic drift between localities. North Head genetic diversity indices estimated between 2002 to 2018 showed relatively constant levels of allelic richness (1.90–2.00) and observed heterozygosity (<i>H</i><sub>O</sub> = 0.231–0.310) along with minor levels of inbreeding (<i>F</i><sub>IS</sub> 0.020–0.052). The identification of some individuals sampled on North Head that were assigned to other populations suggests some sporadic geneflow into the population has occurred and may have assisted with maintaining genetic diversity.</p> <p>These data were used to suggest that the North Head population is distinct from other northern Sydney populations and has relatively constant levels of genetic diversity.</p>
Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements
Open the record for dataset details and reuse information.
50K SNP genotypes of Cameroon Blackbelly and Barbados Blackbelly sheep
Open the record for dataset details and reuse information.
North Pacific harbor porpoise SNP and microhaplotype genotypes, mitochondrial control region haplotype sequences
Open the record for dataset details and reuse information.
SNP genotyping of Lord Howe woodhen (Hypotaenidia sylvestris) from museum skins and contemporary blood samples
Open the record for dataset details and reuse information.
Data from: Comparing methods for SNP calling from Genotyping-By-Sequencing (GBS) data for a large-genome conifer without a published genome sequence
Open the record for dataset details and reuse information.
Data from: SNP genotyping identifies new signatures of selection in a deep sample of West African P. falciparum malaria parasites
Open the record for dataset details and reuse information.
Data from: A high density SNP chip for genotyping great tit (Parus major) populations and its application to studying the genetic architecture of exploration behaviour
Open the record for dataset details and reuse information.
Data from: Multiplex preamplification PCR and microsatellite validation allows accurate single nucleotide polymorphism (SNP) genotyping of historical fish scales
Open the record for dataset details and reuse information.
Data from: SNP genotyping elucidates the genetic diversity of Magna Graecia grapevine germplasm and its historical origin and dissemination
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.