Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38
datasets available to search
ShareScore release 0.9.0
Dataset results
38 results for “single nucleotide polymorphism (SNP)”
LA1141 × OH8245 inbred backcross (IBC) single nucleotide polymorphism (SNP) markers for genetic studies
<p>The LA1141 × OH8245 157 polymorphic SNP markers from an optimized tomato panel Sim et al., 2012 were used for linkage map construction in the BC<sub>2</sub>S<sub>3</sub> IBC and composite interval mapping QTL analysis. Genetic map position and physical position corresponding to Sl4.0 (Hosmani et al., 2019), and flanking sequences are provided.</p>
Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams
<p>DNA were derived from fin tissue samples taken from individual trout captured from Little Jacks Creek, Big Jacks Creek , and Duncan Creek of the Owyhee mountains and Keithly Creek and Upper Mann Creek in the Hitt mountains of Western Idaho, United States. Fin tissues were collected from individual trout from each stream during monthly sampling events in June through October 2020. </p> <p><em>DNA Extraction:</em> Extraction of DNA from caudal fin tissues were performed using Quick-DNA Miniprep Plus purification kits (Zymo Research Inc.©). Small sections of fin tissue (≤ 25 mg) were collected from each sample. This was mixed with a digesting solution comprised of ultra-pure water, solid tissue buffer (Zymo Research Inc.©) and proteinase K. All tissues were digested in sealed microcentrifuge tubes for at minimum 3 h at 55°C in a water bath. We then aliquoted 100 µL of digestion supernatant and combined with 200 µL of genomic binding buffer (Zymo Research Inc.©). DNA was eluted in 50, 75, and 100 µL of elution buffer to determine which volume provided sufficient DNA concentration for genotyping. After it was determined all quantities produced suitable concentrations, going forward, 50 µL of elution buffer used.</p> <p><em>Genotyping:</em> Following extraction, genotyping-in-thousands sequencing took place at the Hagerman National Fish Hatchery’s genetics research facility with the assistance of the Columbia River Intertribal Fish Commission (CRTFC). Genotyping protocols were as described in Campbell et al. (2015) and summarized below. First, samples were prepared for amplification via PCR by combining DNA extracts with a Qiagen Plus multiplex master mix and a species-specific pooled primer mix. This step added the Illumina sequencing primer sites to amplicons. Following the creation of the PCR cocktail, thermocycling was conducted for amplification. Amplified samples were then diluted 20-fold. Diluted samples were transferred to new 96-well PCR plates where two genetic indexes and barcodes provides a unique set of tagging primers to each well and plate. Tagged plates then underwent a second PCR step. After the second PCR, all DNA were transferred to Charm Biotech normalization plates where DNA was bound to wells, washed, and finally eluted. After normalization, all DNA was pooled together and a purification step using magnetized beads in two steps to selectively remove fragments of DNA that are both too large and too small for sequencing. Following purification, each plate was quantified via qPCR using Life Technologies QuantStudio 6 Flex Instrument (Life Technologies). Finally, sequencing was performed using an Illumina HiSeq 1500 instrument.</p> <p><strong>Ancillary peer-reviewed manuscripts:</strong><br> <em>Genotyping protocols</em><br> Campbell NR, Harmon SA, Narum SR. 2015. Genotyping-in-Thousands by sequencing (GT-seq): A cost effective SNP genotyping method based on custom amplicon sequencing. Mol Ecol Resour, 15: 855-867. https://doi.org/10.1111/1755-0998.12357<br> <em>SNP loci reference</em><br> Collins EE, Hargrove JS, Delomas TA, Narum SR. 2020. Distribution of genetic variation underlying adult migration timing in steelhead of the Columbia River basin. Ecology and Evolution, 10(17): 9486-9502. https://doi.org/10.1002/ece3.6641 </p> <p><strong>Data Use</strong>:<br> <em>License</em>: <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY 4.0</a> <br> <em>Recommended Citation</em>: Wooding AP, Narum SR, Pradhan DS. 2022. Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams (0.1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7055582</p> <p>Funding for this project is provided by US National Science Foundation and Idaho EPSCoR through award: OIA-1757324 </p>
Single nucleotide Polymorphism (SNP) identification, genetic diversity, and population structure of Ryegrass from the northeastern highlands of Peru
Open the record for dataset details and reuse information.
Single nucleotide polymorphism (SNP) genotypes of Cashmere goat (Capra hircus) populations from Mongolia
Open the record for dataset details and reuse information.
Data from: Identifying litchi (Litchi chinensis Sonn.) cultivars and their genetic relationships using single nucleotide polymorphism (SNP) markers
Litchi is an important fruit tree in tropical and subtropical areas of the world. However, there is widespread confusion regarding litchi cultivar nomenclature and detailed information of genetic relationships among litchi germplasm is unclear. In the present study, the potential of single nucleotide polymorphism (SNP) for the identification of 96 representative litchi accessions and their genetic relationships in China was evaluated using 155 SNPs that were evenly spaced across litchi genome. Ninety SNPs with minor allele frequencies above 0.05 and a good genotyping success rate were used for further analysis. A relatively high level of genetic variation was observed among litchi accessions, as quantified by the expected heterozygosity (He = 0.305). The SNP based multilocus matching identified two synonymous groups, 'Heiye' and 'Wuye', and 'Chengtuo' and 'Baitangli 1'. A subset of 14 SNPs was sufficient to distinguish all the non-redundant litchi genotypes, and these SNPs were proven to be highly stable by repeated analyses of a selected group of cultivars. Unweighted pair-group method of arithmetic averages (UPGMA) cluster analysis divided the litchi accessions analyzed into four main groups, which corresponded to the traits of extremely early-maturing, early-maturing, middle-maturing, and late-maturing, indicating that the fruit maturation period should be considered as the primary criterion for litchi taxonomy. Two subpopulations were detected among litchi accessions by STRUCTURE analysis, and accessions with extremely early- and late-maturing traits showed membership coefficients above 0.99 for Cluster 1 and Cluster 2, respectively. Accessions with early- and middle-maturing traits were identified as admixture forms with varying levels of membership shared between the two clusters, indicating their hybrid origin during litchi domestication. The results of this study will benefit litchi germplasm conservation programs and facilitate maximum genetic gains in litchi breeding programs.
Data from: Accuracy of assignment of Atlantic salmon (Salmo salar L.) to rivers and regions in Scotland and northeast England based on single nucleotide polymorphism (SNP) markers.
Understanding the habitat use patterns of migratory fish, such as Atlantic salmon (Salmo salar L.), and the natural and anthropogenic impacts on them, is aided by the ability to identify individuals to their stock of origin. Presented here are the results of an analysis of informative single nucleotide polymorphic (SNP) markers for detecting genetic structuring in Atlantic salmon in Scotland and NE England and their ability to allow accurate genetic stock identification. 3,787 fish from 147 sites covering 27 rivers were screened at 5,568 SNP markers. In order to identify a cost-effective subset of SNPs, they were ranked according to their ability to differentiate between fish from different rivers. A panel of 288 SNPs was used to examine both individual assignments and mixed stock fisheries and eighteen assignment units were defined. The results improved greatly on previously available methods and, for the first time, fish caught in the marine environment can be confidently assigned to geographically coherent units within Scotland and NE England, including individual rivers. As such, this SNP panel has the potential to aid understanding of the various influences acting upon Atlantic salmon on their marine migrations, be they natural environmental variations and/or anthropogenic impacts, such as mixed stock fisheries and interactions with marine power generation installations.
Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing
This study aimed to develop a set of SNP markers with high resolution and accuracy within the African buffalo. Such a set can be used, among others, to depict subtle population genetic structure for a better understanding of buffalo population dynamics. In total, 18.5 million DNA sequences of 76 bp were generated by next generation sequencing on an Illumina Genome Analyzer II from a reduced representation library using DNA from a panel of 13 African buffalo representative of the four subspecies. We identified 2534 SNPs with high confidence within the panel by aligning the short sequences to the cattle genome (Bos taurus). The average sequencing depth of the complete aligned set of reads was estimated at 5x, and at 13x when only considering the final set of putative SNPs that passed the filtering criterion. Our set of SNPs was validated by PCR amplification and Sanger sequencing of 15 SNPs. Of these 15 SNPs, 14 amplified successfully and 13 were shown to be polymorphic (success rate: 87%). The fidelity of the identified set of SNPs and potential future applications are finally discussed.
Data from: Multiplex preamplification PCR and microsatellite validation allows accurate single nucleotide polymorphism (SNP) genotyping of historical fish scales
Incorporating historical tissues into the study of ecological, conservation, and management questions can broaden the scope of population genetic research by enhancing our understanding of evolutionary processes and anthropogenic influences on natural populations. Genotyping historical and low-quality samples has been plagued by challenges associated with low amounts of template DNA and the potential for preexisting DNA contamination among samples. We describe a two-step process designed to (i) accurately genotype large numbers of historical low-quality scale samples in a high-throughput format and (ii) screen samples for preexisting DNA contamination. First, we describe how an efficient multiplex preamplification PCR of 45 single nucleotide polymorphisms (SNPs) can generate highly accurate genotypes with low failure and error rates in subsequent SNP genotyping reactions of individual historical scales from sockeye salmon (Oncorhynchus nerka). Second, we demonstrate how the method can be modified for the amplification of microsatellite loci to detect preexisting DNA contamination. A total of 760 individual historical scale and 182 contemporary fin clip samples were genotyped and screened for contamination. Genotyping failure and error rates were exceedingly low and similar for both historical and contemporary samples. Preexisting contamination in 21% of the historical samples was successfully identified by screening the amplified microsatellite loci. The potential for automation, low failure and error rates, and ability to multiplex both the preamplification and subsequent genotyping reactions combine to make the protocol ideally suited for efficiently genotyping large numbers of potentially contaminated low-quality sources of DNA.
Western redcedar single nucleotide polymorphism (SNP) genotyping data for genomic selection and population genetics
<p>Western redcedar (<em>Thuja plicata</em>) Single Nucleotide Polymorphism (SNP) data in Variant Call Format (VCF) for genomic selection training and target populations, genomic selection parents, and self-fertilized (selfing) lines, comprising 4,833 trees.</p> <p>Targeted sequencing-based genotyping was done by Capture-Seq methodology at Rapid Genomics (Neves est al. 2013). A set of 57,000 probes as designed for initial marker discovery, from which a panel of 20,858 probes was selected for genotyping. A set of transcriptomes (Shalev et al. 2018) (PRJNA704616) was aligned to the reference genome to identify SNPs. Candidate probes (120 nt) were initially designed in silico and 57,000 selected by removing candidates with poor base composition for hybridization (GC content <0.2 and >0.6, high G content >0.2 and long homopolymers >7), followed by removing probes aligning to more than one position on the reference genome (≥90% identity and length). The 57,000 probes represent 14,517 scaffolds (average 3.9 probes/scaffold), with 37,275 targeting at least one SNP and 19,725 mapping to intergenic regions not containing pre-identified SNPs. A set of 128 individuals were selected to validate the 57,000 probe panel and associated polymorphisms. Genomic DNA (0.5 ug) was fragmented (mean size 300 bp), followed by repair of ends, phosphorylation, adenylation, ligation of Illumina compatible adapters containing 8bp indexes and 5’ T-overhang, and 10 cycles PCR amplification with universal primers to produce sequencing-ready libraries. Libraries were quantified using PicoGreen. Libraries from 16 samples were pooled, hybridized to the 120 nt RNA probes following Agilent’s SureSelect Target Enrichment System (Agilent Technologies) and sequenced on an Illumina HiSeq X machine with paired-end 150bp cycle for an average sequencing depth per sample of 15X. Sequence data were aligned to the reference genome with BWA-MEM (http://arxiv.org/abs/1303.3997) and sets of four samples were combined to increase sequencing depth for identifying markers. Putative SNPs were identified using Freebayes (http://arxiv.org/abs/1207.3907) in 150bp on either side of the 57,000 probes and filtered probes that had more than 17 SNPs per 420 bp target region (150bp + 120bp + 150bp). The sequencing depth of the probes was used to select the final set of 20,885 probes, removing probes on both sides of the distribution (low and high sequencing depth), for Capture-Seq on the remainder of the samples.</p>
Data from: Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks
An Illumina Infinium SNP genotyping array was constructed for European white oaks. Six individuals of Quercus petraea and Q. robur were considered for SNP discovery using both previously obtained Sanger sequences across 676 gene regions (1371 in vitro SNPs) and Roche 454 technology sequences from 5112 contigs (6542 putative in silico SNPs). The 7913 SNPs were genotyped across the six parental individuals, full-sib progenies (one within each species and two interspecific crosses between Q. petraea and Q. robur) and three natural populations from south-western France that included two additional interfertile white oak species (Q. pubescens and Q. pyrenaica). The genotyping success rate in mapping populations was 80.4% overall and 72.4% for polymorphic SNPs. In natural populations, these figures were lower (54.8% and 51.9%, respectively). Illumina genotype clusters with compression (shift of clusters on the normalized x-axis) were detected in ~25% of the successfully genotyped SNPs and may be due to the presence of paralogues. Compressed clusters were significantly more frequent for SNPs showing a priori incorrect Illumina genotypes, suggesting that they should be considered with caution or discarded. Altogether, these results show a high experimental error rate for the Infinium array (between 15% and 20% of SNPs potentially unreliable and 10% when excluding all compressed clusters), and recommendations are proposed when applying this type of high-throughput technique. Finally, results on diversity levels and shared polymorphisms across targeted white oaks and more distant species of the Quercus genus are discussed, and perspectives for future comparative studies are proposed.
Immunological Change and the Single Nucleotide Polymorphism (SNP) in Children With Irritable Bowel Syndrome (IBS)
ClinicalTrials.gov study NCT01131442. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Data from: Accuracy of assignment of Atlantic salmon (Salmo salar L.) to rivers and regions in Scotland and northeast England based on single nucleotide polymorphism (SNP) markers.
Open the record for dataset details and reuse information.
Data from: Multiplex preamplification PCR and microsatellite validation allows accurate single nucleotide polymorphism (SNP) genotyping of historical fish scales
Open the record for dataset details and reuse information.
Data from: Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks
Open the record for dataset details and reuse information.
Data from: Identifying litchi (Litchi chinensis Sonn.) cultivars and their genetic relationships using single nucleotide polymorphism (SNP) markers
Open the record for dataset details and reuse information.
Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing
Open the record for dataset details and reuse information.
Data from: A second-generation diagnostic single nucleotide polymorphism (SNP)-based assay, optimized to distinguish among eight poplar (Populus L.) species and their early hybrids
Open the record for dataset details and reuse information.
Single Nucleotide Polymorphism (SNP) and Antibody-based Cell Sorting (SNACS): A tool for demultiplexing single-cell DNA sequencing data
GEO Series GSE255224. Homo sapiens. 15 samples. Type: Other.
Copy number variant (CNV) and Single nucleotide polymorphism (SNP) of UCLA hESC lines
GEO Series GSE91072. Homo sapiens. 15 samples. Type: Genome variation profiling by SNP array.
Data from: Genetic diversity among INERA maize inbred lines with single nucleotide polymorphism (SNP) markers and their relationship with CIMMYT, IITA, and temperate lines
Background: Genetic diversity provides the capacity for plants to meet changing environments. It is fundamentally important in crop improvement. Fifty-nine local maize lines developed at INERA and 41 exotic (temperate and tropical) inbred lines were characterized using 1057 SNP markers to (1) analyse the genetic diversity in a diverse set of maize inbred lines; (2) determine the level of genetic diversity in INERA inbred lines and patterns of relationships of these inbred lines developed from two sources; and (3) examine the genetic differences between local and exotic germplasms. Results: Roger's genetic distance for about 64% of the pairs of lines fell between 0.300 and 0.400. Sixty one per cent of the pairs of lines also showed relative kinship values of zero. Model-based population structure analysis and principal component analysis revealed the presence of 5 groups that agree, to some extent, with the origin of the germplasm. There was genetic diversity among INERA inbred lines, which were genetically less closely related and showed a low level of heterozygosity. These lines could be divided into 3 major distinct groups and a mixed group consistent with the source population of the lines. Pairwise comparisons between local and exotic germplasms showed that the temperate and some IITA lines were differentiated from INERA lines. There appeared to be substantial levels of genetic variation between local and exotic germplasms as revealed by missing and unique alleles. Conclusions: Allelic frequency differences observed between the germplasms, together with unique alleles identified within each germplasm, shows the potential for a mutual improvement between the sets of germplasm. The results from this study will be useful to breeders in designing inbred-hybrid breeding programs, association mapping population studies and marker assisted breeding.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.