Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
55
datasets available to search
ShareScore release 0.9.0
Dataset results
55 results for “SNP genotype data”
SNP and indel discovery and genotyping in next-generation sequencing data
<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels >50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Data from: Phylogenetic relationships, breeding implications, and cultivation history of Hawaiian taro (Colocasia esculenta) through genome-wide SNP genotyping
Taro, Colocasia esculenta, is one of the world's oldest root crops and of particular economic and cultural significance in Hawai'i, where historically more than 150 different landraces were grown. We developed a genome-wide set of more than 2400 high-quality single nucleotide polymorphism (SNP) markers from 70 taro accessions of Hawaiian, South Pacific, Palauan, and mainland Asian origins, with several objectives: (a) uncover the phylogenetic relationships between Hawaiian and other Pacific landraces, (b) shed light on the history of taro cultivation in Hawai'i, and (c) develop a tool to discriminate among Hawaiian and other taros. We found that almost all existing Hawaiian landraces fall into five monophyletic groups that are largely consistent with the traditional Hawaiian classification based on morphological characters, e.g., leaf shape and petiole color. Genetic diversity was low within these clades but considerably higher between them. Population structure analyses further indicated that the diversification of taro in Hawai'i most likely occurred by a combination of frequent somatic mutation and occasional hybridization. Unexpectedly, the South Pacific accessions were found nested within the clades mainly composed of Hawaiian accessions, rather than paraphyletic to them. This suggests that the origin of clades identified here preceded the colonization of Hawai'i, and that early Polynesian settlers brought taro landraces from different clades with them. In the absence of a sequenced genome, this marker set provides a valuable resource towards obtaining a genetic linkage map, and to study the genetic basis of phenotypic traits of interest to taro breeding such as disease resistance.
Mingrelian SNP Genotype Data
<p>This dataset contains data from 645,337 single nucleotide polymorphisms (SNPs) that were genotyped on GenoChip 2+ microarrays. The SNP data were ascertained from the mtDNA, Y-chromosome and autosomes for each individual, depending on their biological sex. In total, 5,205 mtDNA and 10,272 Y-chromosome SNPs were extracted from the array data. These data files have been uploaded as .csv files and also be uploaded as plink-formatted files. Details about the analysis of the SNP data can be found in the associated manuscript:</p><p>Theodore G Schurr, Ramaz Shengelia, Michel Shamoon-Pour, David Chitanava, Shorena Laliashvili, Irma Laliashvili, Redate Kibret, Yanu Kume-Kangkolo, Irakli Akhvlediani, Lia Bitadze, Iain Mathieson, Aram Yardumian, Genetic Analysis of Mingrelians Reveals Long-Term Continuity of Populations in Western Georgia (Caucasus), <i>Genome Biology and Evolution</i>, 2023; evad198, <a href="https://doi.org/10.1093/gbe/evad198">https://doi.org/10.1093/gbe/evad198</a></p>
Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs
<p>The development of high-throughput sequencing has prompted a transition in wildlife genetics from using microsatellites toward sets of Single Nucleotide Polymorphisms (SNPs). However, genotyping large numbers of targeted SNPs using non-invasive samples remains challenging due to relatively large DNA input requirements. Recently, target enrichment has emerged as a promising approach requiring little template DNA. We assessed the efficacy of Tecan Genomics' Allegro Targeted Genotyping (ATG) for generating genome-wide SNP data in feral horses using DNA isolated from fecal swabs. Total and host-specific DNA were quantified for 989 samples collected as part of a long-term individual-based study of feral horses on Sable Island, Nova Scotia, Canada, using dsDNA fluorescence and a host-specific qPCR assay, respectively. Forty-eight samples representing 44 individuals containing at least 10ng of host DNA (ATG's recommended minimum input) were genotyped using a custom multiplex panel targeting 279 SNPs. Genotyping accuracy and consistency were assessed by contrasting ATG genotypes with those obtained from the same individuals with SNP microarrays, and from multiple samples from the same horse, respectively. 62% of swabs yielded the minimum recommended amount of host DNA for ATG. Ignoring samples that failed to amplify, ATG recovered an average of 86.7% targeted sites per sample, while genotype concordance between ATG and SNP microarrays was 98.5%. The repeatability of genotypes from the same individual approached unity with an average of 99.9%. This study demonstrates the suitability of ATG for genome-wide, non-invasive targeted SNP genotyping, and will facilitate further ecological and conservation genetics research in equids and related species.</p>
Autosomal SNP-genotype data of brown bears (Ursus arctos) in Finland
<p>Harmonising methodology between countries is crucial in transborder population monitoring. However, immediate application of alleged, established DNA-based methods across the extended area can entail drawbacks and may lead to biases. Therefore, genetic methods need to be tested across the whole area before being deployed. Around 4,500 brown bears (<em>Ursus arctos</em>) live in Norway, Sweden, and Finland and they are divided into the western (Scandinavian) and eastern (Karelian) population. Both populations have recovered and are connected via asymmetric migration. DNA-based population monitoring in Norway and Sweden uses the same set of genetic markers. With Finland aiming to implement monitoring, we tested the available SNP-panel developed to assess brown bears in Norway and Sweden, on tissue samples from a representative set of 93 legally harvested individuals from Finland. The aim was to test for ascertainment bias and evaluate its suitability for DNA-based transnational-monitoring covering all three countries. We compared results to the performance of microsatellite genotypes of the same individuals in Finland and against SNP-genotypes from individuals sampled in Sweden (<em>N</em>=95) and Norway (<em>N</em>=27). In Finland, a higher resolution for individual identification was obtained for SNPs (PI=1.18E-27) compared to microsatellites (PI=4.2E-11). Compared to Norway and Sweden, probability of identity of the SNP-panel was slightly higher and expected heterozygosity lower in Finland indicating ascertainment bias. Yet, our evaluation show that the available SNP-panel outperforms the microsatellite panel currently applied in Norway and Sweden. The SNP-panel represents a powerful tool that could aid improving transnational DNA-based monitoring of brown bears across these three countries.</p>
Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs
Open the record for dataset details and reuse information.
SNP genotype and hyperspectral reflectance data from: Ensembles of genomic and hyperspectral imaging-based prediction enable selection for reduced deoxynivalenol content in wheat grains
Open the record for dataset details and reuse information.
Data from: Phylogenetic relationships, breeding implications, and cultivation history of Hawaiian taro (Colocasia esculenta) through genome-wide SNP genotyping
Open the record for dataset details and reuse information.
Autosomal SNP-genotype data of brown bears (Ursus arctos) in Finland
Open the record for dataset details and reuse information.
Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates
Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage >30, but remained >0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from >5 to >30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.
Data from: Phylogeography and adaptation genetics of stickleback from the Haida Gwaii archipelago revealed using genome-wide SNP genotyping
Threespine stickleback populations are model systems for studying adaptive evolution and the underlying genetics. In lakes on the Haida Gwaii archipelago (off western Canada), stickleback have undergone a remarkable local radiation and show phenotypic diversity matching that seen throughout the species distribution. To provide a historical context for this radiation, we surveyed genetic variation at >1000 single nucleotide polymorphism (SNP) loci in stickleback from over 100 populations. SNPs included markers evenly distributed throughout genome and candidate SNPs tagging adaptive genomic regions. Based on evenly distributed SNPs, the phylogeographic pattern differs substantially from the disjunct pattern previously observed between two highly divergent mtDNA lineages. The SNP tree instead shows extensive within watershed population clustering and different watersheds separated by short branches deep in the tree. These data are consistent with separate colonizations of most watersheds, despite underlying genetic connections between some independent drainages. This supports previous suppositions that morphological diversity observed between watersheds has been shaped independently, with populations exhibiting complete loss of lateral plates and giant size each occurring in several distinct clades. Throughout the archipelago, we see repeated selection of SNPs tagging candidate freshwater adaptive variants at several genomic regions differentiated between marine–freshwater populations on a global scale (e.g. EDA, Na/K ATPase). In estuarine sites, both marine and freshwater allelic variants were commonly detected. We also found typically marine alleles present in a few freshwater lakes, especially those with completely plated morphology. These results provide a general model for postglacial colonization of freshwater habitat by sticklebacks and illustrate the tremendous potential of genome-wide SNP data sets hold for resolving patterns and processes underlying recent adaptive divergences.
Data from: SNP genotyping identifies new signatures of selection in a deep sample of West African P. falciparum malaria parasites
We used a high density SNP array to genotype 75 P. falciparum isolates recently collected from Senegal and The Gambia in order to search for signals of selection in this malaria endemic region. We found little geographic or temporal stratification of the genetic diversity among the sampled parasites. Through application of the iHS and REHH haplotype-based tests for positive selection, we found evidence of recent selective sweeps at a known drug resistance locus, at several known antigenic loci, and at several genomic regions not previously identified as sites of recent selection. We discuss the value of deep population-specific genomic analyses for identifying selection signals within sampled endemic populations of parasites, which may correspond to local selection pressures such as distinctive therapeutic regimes or mosquito vectors.
Data from: A high density SNP chip for genotyping great tit (Parus major) populations and its application to studying the genetic architecture of exploration behaviour
High density SNP microarrays ('SNP chips') are a rapid, accurate and efficient method for genotyping several hundred thousand polymorphisms in large numbers of individuals. While SNP chips are routinely used in human genetics and in animal and plant breeding, they are less widely used in evolutionary and ecological research. In this paper we describe the development and application of a high density Affymetrix Axiom chip with around 500 000 SNPs, designed to perform genomics studies of great tit (Parus major) populations. We demonstrate that the per-SNP genotype error rate is well below 1% and that the chip can also be used to identify structural or copy number variation (CNVs). The chip is used to explore the genetic architecture of exploration behaviour (EB), a personality trait that has been widely studied in great tits and other species. No SNPs reached genome-wide significance, including at DRD4, a candidate gene. However, EB is heritable and appears to have a polygenic architecture. Researchers developing similar SNP chips may note: (i) SNPs previously typed on alternative platforms are more likely to be converted to working assays, (ii) detecting SNPs by more than one pipeline, and in independent datasets, ensures a high proportion of working assays, (iii) allele frequency ascertainment bias is minimised by performing SNP discovery in individuals from multiple populations and (iv) samples with the lowest call rates tend to also have the greatest genotyping error rates.
Data from: Multiplex preamplification PCR and microsatellite validation allows accurate single nucleotide polymorphism (SNP) genotyping of historical fish scales
Incorporating historical tissues into the study of ecological, conservation, and management questions can broaden the scope of population genetic research by enhancing our understanding of evolutionary processes and anthropogenic influences on natural populations. Genotyping historical and low-quality samples has been plagued by challenges associated with low amounts of template DNA and the potential for preexisting DNA contamination among samples. We describe a two-step process designed to (i) accurately genotype large numbers of historical low-quality scale samples in a high-throughput format and (ii) screen samples for preexisting DNA contamination. First, we describe how an efficient multiplex preamplification PCR of 45 single nucleotide polymorphisms (SNPs) can generate highly accurate genotypes with low failure and error rates in subsequent SNP genotyping reactions of individual historical scales from sockeye salmon (Oncorhynchus nerka). Second, we demonstrate how the method can be modified for the amplification of microsatellite loci to detect preexisting DNA contamination. A total of 760 individual historical scale and 182 contemporary fin clip samples were genotyped and screened for contamination. Genotyping failure and error rates were exceedingly low and similar for both historical and contemporary samples. Preexisting contamination in 21% of the historical samples was successfully identified by screening the amplified microsatellite loci. The potential for automation, low failure and error rates, and ability to multiplex both the preamplification and subsequent genotyping reactions combine to make the protocol ideally suited for efficiently genotyping large numbers of potentially contaminated low-quality sources of DNA.
Impatiens glandulifera SNP and SilicoDArT genotyping data
<p>We conducted genomic characterization based on SNP and SilicoDArT markers on the invasive Himalayan balsam (<i>Impatiens glandulifera</i>) plants originating from the native and non-native regions of their distribution. When genetic relationships were explored by PCoA based on SNP and SilicoDArT marker data, the first, second and third principal coordinates explained altogether 37.4% and 31.0% of the variability, respectively. Samples from the UK, Canada and Pakistan grouped together, while Indian plants were clearly distinct based on SNP markers but relatively close to the UK-Canada-Pakistan group based on SilicoDArT markers. Constructed trees differentiated the individuals into clusters resembling the patterns observed by PCoA.<span> The Bayesian BAPS analysis revealed that the individuals were distributed in seven clusters, representing samples from each of the four Finnish populations, India, Pakistan and the combination of the UK and Canada. Similar clustering was visible in the constructed UPGMA tree. The Indian cluster did not display any ancestral gene flow with the others, while the Pakistani cluster showed ancestral gene flow only with the combined UK and Canada cluster. Furthermore, the latter cluster displayed ancestral gene flow with the Finnish populations varying from 0% to 3.1%. The AMOVA analysis showed that 45% and 26% of genetic variation was present among the <i>I. glandulifera</i> groups/populations and the rest within them based on SNP and SilicoDArT markers, respectively. Overall, the Bayesian BAPS analysis</span> <span>and the following gene flow network were the most informative tools for resolving relationships among native and introduced plants. </span></p>
Western redcedar single nucleotide polymorphism (SNP) genotyping data for genomic selection and population genetics
<p>Western redcedar (<em>Thuja plicata</em>) Single Nucleotide Polymorphism (SNP) data in Variant Call Format (VCF) for genomic selection training and target populations, genomic selection parents, and self-fertilized (selfing) lines, comprising 4,833 trees.</p> <p>Targeted sequencing-based genotyping was done by Capture-Seq methodology at Rapid Genomics (Neves est al. 2013). A set of 57,000 probes as designed for initial marker discovery, from which a panel of 20,858 probes was selected for genotyping. A set of transcriptomes (Shalev et al. 2018) (PRJNA704616) was aligned to the reference genome to identify SNPs. Candidate probes (120 nt) were initially designed in silico and 57,000 selected by removing candidates with poor base composition for hybridization (GC content <0.2 and >0.6, high G content >0.2 and long homopolymers >7), followed by removing probes aligning to more than one position on the reference genome (≥90% identity and length). The 57,000 probes represent 14,517 scaffolds (average 3.9 probes/scaffold), with 37,275 targeting at least one SNP and 19,725 mapping to intergenic regions not containing pre-identified SNPs. A set of 128 individuals were selected to validate the 57,000 probe panel and associated polymorphisms. Genomic DNA (0.5 ug) was fragmented (mean size 300 bp), followed by repair of ends, phosphorylation, adenylation, ligation of Illumina compatible adapters containing 8bp indexes and 5’ T-overhang, and 10 cycles PCR amplification with universal primers to produce sequencing-ready libraries. Libraries were quantified using PicoGreen. Libraries from 16 samples were pooled, hybridized to the 120 nt RNA probes following Agilent’s SureSelect Target Enrichment System (Agilent Technologies) and sequenced on an Illumina HiSeq X machine with paired-end 150bp cycle for an average sequencing depth per sample of 15X. Sequence data were aligned to the reference genome with BWA-MEM (http://arxiv.org/abs/1303.3997) and sets of four samples were combined to increase sequencing depth for identifying markers. Putative SNPs were identified using Freebayes (http://arxiv.org/abs/1207.3907) in 150bp on either side of the 57,000 probes and filtered probes that had more than 17 SNPs per 420 bp target region (150bp + 120bp + 150bp). The sequencing depth of the probes was used to select the final set of 20,885 probes, removing probes on both sides of the distribution (low and high sequencing depth), for Capture-Seq on the remainder of the samples.</p>
Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements
Sample availability limits population genetics research on many species, especially taxa from regions with high diversity. However, many such species are well represented in museum collections assembled before the molecular era. Development of techniques to recover genetic data from these invaluable specimens will benefit biodiversity science. Using a mixture of freshly preserved and historical tissue samples, and a sequence capture probe set targeting >5000 loci, we produced high-confidence genotype calls on thousands of single nucleotide polymorphisms (SNPs) in each of five South-East Asian bird species and their close relatives (N = 27–43). On average, 66.2% of the reads mapped to the pseudo-reference genome of each species. Of these mapped reads, an average of 52.7% was identified as PCR or optical duplicates. We achieved deeper effective sequencing for historical samples (122.7×) compared to modern samples (23.5×). The number of nucleotide sites with at least 8× sequencing depth was high, with averages ranging from 0.89 × 106 bp (Arachnothera, modern samples) to 1.98 × 106 bp (Stachyris, modern samples). Linear regression revealed that the amount of sequence data obtained from each historical sample (represented by per cent of the pseudo-reference genome recovered with ≥8× sequencing depth) was positively and significantly (P ≤ 0.013) related to how recently the sample was collected. We observed characteristic post-mortem damage in the DNA of historical samples. However, we were able to reduce the error rate significantly by truncating ends of reads during read mapping (local alignment) and conducting stringent SNP and genotype filtering.
Data from: High-throughput SNP genotyping of historical and modern samples of five bird species via sequence capture of ultraconserved elements
Open the record for dataset details and reuse information.
Data from: Comparing methods for SNP calling from Genotyping-By-Sequencing (GBS) data for a large-genome conifer without a published genome sequence
Open the record for dataset details and reuse information.
Data from: SNP genotyping identifies new signatures of selection in a deep sample of West African P. falciparum malaria parasites
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.