Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,019
datasets available to search
ShareScore release 0.9.0
Dataset results
1,019 results for “SNP”
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S1 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S0 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Data for "How Array Design Creates SNP Ascertainement Bias"
<p>The repository contains the raw SNP data in vcf format for the publication "How Array Design Creates SNP Ascertainment Bias". Note that the variants are <strong>not</strong> filtered at this timepoint. Samples starting with pl_ are pooled sequences of ~10 individuals while samples starting with i_ were individually sequenced. Please find detailed information about samples, raw sequencing data and SNP calling pipeline in the linked preprint (<a href="https://doi.org/10.1101/833541">https://doi.org/10.1101/833541</a>)/ publication (<a href="https://doi.org/10.1371/journal.pone.0245178">https://doi.org/10.1371/journal.pone.0245178</a>). In case you need additional information, please contact<a href="mailto:johannes.geibel@uni-goettingen.de"> johannes.geibel@uni-goettingen.de</a></p>
Genotyping of the Chinese Spring x Renan mapping population with the TaBW280K SNP array
<p>The TaBW280K SNP array (Rimbert et al., PLoS ONE 2018) was used to genotype 430 Single Seed Descent (SSD) individuals<br> derived from a cross between Chinese Spring and Renan (CsRe; Choulet et al., Science 2014). Out of the 280,226 probesets, 85,276 were found to be polymorphic between the two parental lines and PHR on the population. Eventually, 83,721 (98.2%) SNPs were genetically mapped in 21 linkage groups corresponding to the 21 chromosomes of bread wheat, with no unlinked markers. This file contains the genotyping data of the 430 SSD lines.<br> </p>
SNP and indel discovery and genotyping in next-generation sequencing data
<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels >50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Gene co-ordinates, expression levels; SNP identifiers and functions for Drosophila melanogaster (Sussex LHM population)
<p>Data for SNP context information to add to GWAS results. Specifically, SNP functions, sex-bias in gene expression, official SNP idenfiers from NCBI dbSNP, and gene positions and names (from UCSC Genome Browswer). Most of the input files are on-line and their URLs are stated in the code (make_dmel_accessory_data.sh). Also includes code, logs, and exploratory graphs.</p>
LA1141 × OH8245 inbred backcross (IBC) single nucleotide polymorphism (SNP) markers for genetic studies
<p>The LA1141 × OH8245 157 polymorphic SNP markers from an optimized tomato panel Sim et al., 2012 were used for linkage map construction in the BC<sub>2</sub>S<sub>3</sub> IBC and composite interval mapping QTL analysis. Genetic map position and physical position corresponding to Sl4.0 (Hosmani et al., 2019), and flanking sequences are provided.</p>
Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams
<p>DNA were derived from fin tissue samples taken from individual trout captured from Little Jacks Creek, Big Jacks Creek , and Duncan Creek of the Owyhee mountains and Keithly Creek and Upper Mann Creek in the Hitt mountains of Western Idaho, United States. Fin tissues were collected from individual trout from each stream during monthly sampling events in June through October 2020. </p> <p><em>DNA Extraction:</em> Extraction of DNA from caudal fin tissues were performed using Quick-DNA Miniprep Plus purification kits (Zymo Research Inc.©). Small sections of fin tissue (≤ 25 mg) were collected from each sample. This was mixed with a digesting solution comprised of ultra-pure water, solid tissue buffer (Zymo Research Inc.©) and proteinase K. All tissues were digested in sealed microcentrifuge tubes for at minimum 3 h at 55°C in a water bath. We then aliquoted 100 µL of digestion supernatant and combined with 200 µL of genomic binding buffer (Zymo Research Inc.©). DNA was eluted in 50, 75, and 100 µL of elution buffer to determine which volume provided sufficient DNA concentration for genotyping. After it was determined all quantities produced suitable concentrations, going forward, 50 µL of elution buffer used.</p> <p><em>Genotyping:</em> Following extraction, genotyping-in-thousands sequencing took place at the Hagerman National Fish Hatchery’s genetics research facility with the assistance of the Columbia River Intertribal Fish Commission (CRTFC). Genotyping protocols were as described in Campbell et al. (2015) and summarized below. First, samples were prepared for amplification via PCR by combining DNA extracts with a Qiagen Plus multiplex master mix and a species-specific pooled primer mix. This step added the Illumina sequencing primer sites to amplicons. Following the creation of the PCR cocktail, thermocycling was conducted for amplification. Amplified samples were then diluted 20-fold. Diluted samples were transferred to new 96-well PCR plates where two genetic indexes and barcodes provides a unique set of tagging primers to each well and plate. Tagged plates then underwent a second PCR step. After the second PCR, all DNA were transferred to Charm Biotech normalization plates where DNA was bound to wells, washed, and finally eluted. After normalization, all DNA was pooled together and a purification step using magnetized beads in two steps to selectively remove fragments of DNA that are both too large and too small for sequencing. Following purification, each plate was quantified via qPCR using Life Technologies QuantStudio 6 Flex Instrument (Life Technologies). Finally, sequencing was performed using an Illumina HiSeq 1500 instrument.</p> <p><strong>Ancillary peer-reviewed manuscripts:</strong><br> <em>Genotyping protocols</em><br> Campbell NR, Harmon SA, Narum SR. 2015. Genotyping-in-Thousands by sequencing (GT-seq): A cost effective SNP genotyping method based on custom amplicon sequencing. Mol Ecol Resour, 15: 855-867. https://doi.org/10.1111/1755-0998.12357<br> <em>SNP loci reference</em><br> Collins EE, Hargrove JS, Delomas TA, Narum SR. 2020. Distribution of genetic variation underlying adult migration timing in steelhead of the Columbia River basin. Ecology and Evolution, 10(17): 9486-9502. https://doi.org/10.1002/ece3.6641 </p> <p><strong>Data Use</strong>:<br> <em>License</em>: <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY 4.0</a> <br> <em>Recommended Citation</em>: Wooding AP, Narum SR, Pradhan DS. 2022. Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams (0.1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7055582</p> <p>Funding for this project is provided by US National Science Foundation and Idaho EPSCoR through award: OIA-1757324 </p>
Rye600K SNP Array 'Lo7' Map
<p>Mapping position of rye600K SNP array markers developed by Bauer., et al. 2017, on the reference chromosomal-scale rye reference genome 'Lo7'. Shared in the hope that it provides a valuable and easily-accessible resource for further studies in rye genomics.</p> <p><strong>Method:</strong></p> <p>Positional data of 600K SNP markers was obtained by mapping each of the 600843 SNP marker sequences to the rye reference genome ‘Lo7’ using NCBI blastn (v. 2.9.0+) function. Mapping position of SNPs were hereafter stringently filtered for I) complete SNP sequence alignment, and II) maximum of 1 mismatch to ensure an accurate positioning. </p> <p> </p> <p> </p>
Genetic diversity, population structure, and linkage disequilibrium among tropical quality protein maize (QPM) lines assessed with high-density SNP markers
<p>The study of genetic diversity (GD), population structure, and linkage disequilibrium (LD) provides a better understanding of the genetic relationships between individuals in a population which can be utilized in crop research and improvement. Genotyping-by-sequencing (GBS) was used to detect and genotype single nucleotide polymorphisms (SNPs) in a collection of 74 quality protein maize (QPM) lines and further to characterize their genetic diversity, population structure, and linkage disequilibrium. A total of 235,214 high-quality SNPs were used for different genetic analyses except for structure analysis where 11,950 SNPs were used. Analysis of molecular variance (AMOVA) based on these SNPs revealed high genetic heterozygosity among the five populations with 1% of the total genetic variation present among the subpopulations and 99% of the variation among individuals within the populations. Population structure analysis using Bayesian-based clustering revealed that the 74 lines could be clustered into four groups. However, neighbor-joining trees indicate the lines are grouped into three major clusters. Further analysis using principal component analyses (PCA) clustered the genotypes into five groups which are concordant with the groups based on pedigree information. Higher genetic diversity was detected in population 1 with a GD value of 0.484 and the lowest in population 5 (0.396) and overall, with a mean of 0.434. The LD pattern in the quality protein maize was investigated and we observed a relatively rapid LD decay of 3.53kb and 10.66kb at r<sup>2</sup> =0.2 and r<sup>2</sup>= 0.1, respectively. Our findings provide important information for future Linkage mapping studies, genome-wide association analyses, and marker-assisted selective breeding of maize as well as genomic prediction-based selection in tropical germplasm.</p>
High density SNP genotypes (Infinium Human CytoSNP-850K v1.2 BeadChip) of ERAP2-WT and ERAP2-KO Birdshot LCL
<p>SNP genotype data was performed on DNA isolated from WT and CRISPR-Cas9 edited LCLs (ERAP2-KO) according to standard procedures using the Infinium Human CytoSNP-850K v1.2 BeadChip (Illumina, San Diego, CA, USA). SNP-array results and data analysis were carried out using NxClinical software v5.1 (BioDiscovery, Los Angeles, CA, USA). Human genome build Feb. 2009 GRCh37/hg19 was used. </p>
Planform change and Fundulus SNP data for small watersheds in South Mississippi and Louisiana
<p>Fluvial geomorphic processes and the resulting patterns of landform morphogenesis affect the distribution and connectivity of habitat patches for aquatic organisms. Human alterations to fluvial geomorphic processes may affect local habitat quality and stability, and affect connectivity of habitat patches by altering the distribution, supply, and movement of landform-generating materials. This dataset examines 17 watersheds in south Mississippi and southeastern Louisiana and was used in preparation of a manuscript addressing the hypothesis that elevated planform movement, indicative of advanced fluvial erosion, would cause fragmentation among populations of a headwater specialist (Blackspotted Topminnow <em>Fundulus olivaceus</em>). The dataset includes numerous spatial features derived from the NHD+ dataset used in planform measurements, spatial features digitized from NAPP and NAIP aerial imagery measuring planform characteristics and dynamics, additional metrics of each watershed, and a population genetics dataset of single nucleotide polymorphisms (SNPs) for multiple individuals at multiple sites per watershed. Associated code to recreate all analyses in the manuscript is provided.</p>
Anolis carolinensis character displacement SNP
<p>Here are six files that provide details for all 44,120 identified single nucleotide polymorphisms (SNPs) or the 215 outlier SNPs associated with the evolution of rapid character displacement among replicate islands with (2Spp) and without competition (1Spp) between two <em>Anolis</em> species. On 2Spp islands, <em>A. carolinensis</em> occurs higher in trees and have evolved larger toe pads. Among 1Spp and 2Spp island populations, we identify 44,120 SNPs, with 215-outlier SNPs with improbably large F<sub>ST</sub> values, low nucleotide variation, greater linkage than expected, and these SNPs are enriched for animal walking behavior. Thus, we conclude that these 215-outliers are evolving by natural selection in response to the phenotypic convergent evolution of character displacement. There are two, non-mutually exclusive perspective of these nucleotide variants. One is character displacement is convergent: all 215 outlier SNPs are shared among 3 out of 5 2Spp island and 24% of outlier SNPS are shared among all five out of five 2Spp island. Second, character displacement is genetically redundant because the allele frequencies in one or more 2Spp are similar to 1Spp islands: among one or more 2Spp islands 33% of outlier SNPS are within the range of 1Spp MiAF and 76% of outliers are more similar to 1Spp island than mean MiAF of 2Spp islands. Focusing on convergence SNP is scientifically more robust, yet it distracts from the perspective of multiple genetic solutions that enhances the rate and stability of adaptive change.</p> <p>The six files include: a description of eight islands, details of 94 individuals, and four files on SNPs. The four SNP files include the VCF files for 94 individuals with 44KSNPs and two files (Excel sheet/tab-delimited file) with F<sub>ST</sub>, p-values and outlier status for all 44,120 identified single nucleotide polymorphisms (SNPs) associated with the evolution of rapid character displacement. The sixth file is a detailed file on the 215 outlier SNPs.</p> <p>Complete sequence data is available at Bioproject PRJNA833453, which including samples not included in this study. The 94 individuals used in this study are described in "Supplemental_Sample_description.txt"</p>
Development of a high-density 665 K SNP array for rainbow trout genome-wide genotyping. Supplemental VCF file
<p>Single nucleotide polymorphism (SNP) arrays, also named « SNP chips », enable very large numbers of individuals to be genotyped at a targeted set of thousands of genome-wide identified markers. We used preexisting variant datasets from USDA, a French commercial line and 30X-coverage whole genome sequencing of INRAE isogenic lines to develop an Affymetrix 665 K SNP array (HD chip) for rainbow trout. In total, we identified 32,372,492 SNPs that were polymorphic in the USDA or INRAE databases. A subset of identified SNPs were selected for inclusion on the chip, prioritizing SNPs whose flanking sequence uniquely aligned to the Swanson reference genome, with homogenous repartition over the genome and the highest Minimum Allele Frequency in both USDA and French databases. Of the 664,531 SNPs which passed the Affymetrix quality filters and were manufactured on the HD chip, 65.3% and 60.9% passed filtering metrics and were polymorphic in two other distinct French commercial populations in which, respectively, 288 and 175 sampled fish were genotyped. Only 576,118 SNPs mapped uniquely on both Swanson and Arlee reference genomes, and 12,071 SNPs did not map at all on the Arlee reference genome. Among those 576,118 SNPs, 38,948 SNPs were kept from the commercially available medium-density 57K SNP chip. We demonstrate the utility of the HD chip by describing the high rates of linkage disequilibrium at 2 kb to 10 kb in the rainbow trout genome in comparison to the linkage disequilibrium observed at 50 kb to 100 kb which are usual distances between markers of the medium-density chip.</p> <p> </p> <p>File submitted correspond to the supplementary data 1 of the publication (under submission) : INRAE_USDA_MAF1.vcf.gz</p>
Construction of a SNP fingerprinting database and population genetic analysis of 329 cauliflower cultivars
<p>The VCF file contains the information of 1662 SNP sites of 820 cauliflower inbred lines that were filtered according to a series of stringent conditions.</p>
70K SNP array data for Lumpfish (Cyclopterus lumpus) across the trans-Atlantic
<p>In marine species with large populations and high dispersal potential, large-scale genetic differences and clinal trends in allele frequency can provide insight into the evolutionary processes that shape diversity. Lumpfish, <em>Cyclopterus lumpus</em>, is found throughout the North Atlantic and has traditionally been harvested for roe and more recently used as a cleaner fish in salmon aquaculture. We used a 70K SNP array to evaluate trans-Atlantic differentiation, genetic structuring, and clinal variation across the North Atlantic. Basin-scale structuring between the Northeast and Northwest Atlantic was significant, with enrichment for loci associated with developmental/mitochondrial function. We identified a putative structural variant on chromosome 2, likely contributing to differentiation between Northeast and Northwest Atlantic Lumpfish, and consistent with post-glacial trans-Atlantic secondary contact. Redundancy Analysis identified climate associations both in the Northeast (<em>N</em> = 1269 loci) and Northwest (<em>N</em> = 1637 loci), with 103 shared loci between them. Clinal patterns in allele frequencies were observed in some loci (15% - Northwest and 5% - Northeast) of which 708 loci were shared and involved with growth, developmental processes, and locomotion. The combined evidence of trans-Atlantic differentiation, environmental associations, and clinal loci, suggests that both regional and large-scale potentially-adaptive population structuring is present across the North Atlantic.</p>
Fig. 1 in Morphological characters and SNP markers suggest hybridization and introgression in sympatric populations of the pleurocarpous mosses Homalothecium lutescens and H. sericeum
Fig. 1 Measurements of branch leaf characteristics of Homalothecium spp. Leaf characters (L1-L18) are explained in Table 3
Fig. 4 in Morphological characters and SNP markers suggest hybridization and introgression in sympatric populations of the pleurocarpous mosses Homalothecium lutescens and H. sericeum
Fig. 4 Relationship between width and length of leaf lamina of Homalothecium leaf specimens (N = 240) collected from the allopatric populations of H. lutescens (N = 60) and H. sericeum (N = 59) and the
Fig. 6 in Morphological characters and SNP markers suggest hybridization and introgression in sympatric populations of the pleurocarpous mosses Homalothecium lutescens and H. sericeum
Fig. 6 The positions of putatively hybrid sporophytes from the sympatric populations of H. lutescens and H. sericeum superimposed in the PCA from Fig. 5, based on leaf morphology of the maternal gametophytes. The morphospace of leaves from allopatric populations of H. lutescens and H. sericeum are shown as encircled surfaces (blue circle = H. lutescens and yellow circle = H. sericeum). The red circle represents the morphospace of individuals from the sympatric populations. Each square represents a hybrid sporophyte specimen, collected on its
Fig. 5 in Morphological characters and SNP markers suggest hybridization and introgression in sympatric populations of the pleurocarpous mosses Homalothecium lutescens and H. sericeum
Fig. 5 Principal component analysis of 14 leaf characters from 240 specimens representing allopatric and sympatric populations of Homalothecium lutescens and H. sericeum. The first two axes (PC1 and PC2) representing together 36% of variation are shown. The colours and shapes of data points correspond to the population of the specimens. Leaf
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.