Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

210

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

210 results for “single nucleotide polymorphism”

Learn how ShareScore rates datasets ↗
zenodo44/100

LA1141 × OH8245 inbred backcross (IBC) single nucleotide polymorphism (SNP) markers for genetic studies

<p>The LA1141 &times; OH8245 157 polymorphic SNP markers from an optimized tomato panel Sim et al., 2012&nbsp;were used for linkage map construction in the BC<sub>2</sub>S<sub>3</sub>&nbsp;IBC and composite interval mapping QTL analysis. Genetic map position and physical position corresponding to&nbsp;Sl4.0 (Hosmani et al., 2019), and flanking sequences are provided.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams

<p>DNA were derived from fin tissue samples taken from individual trout captured from Little Jacks Creek, Big Jacks Creek , and Duncan Creek of the Owyhee mountains and Keithly Creek and Upper Mann Creek in the Hitt mountains of Western Idaho, United States. Fin tissues were collected from individual trout from each stream during monthly sampling events in June through October 2020.&nbsp;</p> <p><em>DNA Extraction:</em> Extraction of DNA from caudal fin tissues were performed using Quick-DNA Miniprep Plus purification kits (Zymo Research Inc.&copy;). Small sections of fin tissue (&le; 25 mg) were collected from each sample. This was mixed with a digesting solution comprised of ultra-pure water, solid tissue buffer (Zymo Research Inc.&copy;) and proteinase K. All tissues were digested in sealed microcentrifuge tubes for at minimum 3 h at 55&deg;C in a water bath. We then aliquoted 100 &micro;L of digestion supernatant and combined with 200 &micro;L of genomic binding buffer (Zymo Research Inc.&copy;). DNA was eluted in 50, 75, and 100 &micro;L of elution buffer to determine which volume provided sufficient DNA concentration for genotyping. After it was determined all quantities produced suitable concentrations, going forward, 50 &micro;L of elution buffer used.</p> <p><em>Genotyping:</em> Following extraction, genotyping-in-thousands sequencing took place at the Hagerman National Fish Hatchery&rsquo;s genetics research facility with the assistance of the Columbia River Intertribal Fish Commission (CRTFC). Genotyping protocols were as described in Campbell et al. (2015) and summarized below. First, samples were prepared for amplification via PCR by combining DNA extracts with a Qiagen Plus multiplex master mix and a species-specific pooled primer mix. This step added the Illumina sequencing primer sites to amplicons. Following the creation of the PCR cocktail, thermocycling was conducted for amplification. Amplified samples were then diluted 20-fold. Diluted samples were transferred to new 96-well PCR plates where two genetic indexes and barcodes provides a unique set of tagging primers to each well and plate. Tagged plates then underwent a second PCR step. After the second PCR, all DNA were transferred to Charm Biotech normalization plates where DNA was bound to wells, washed, and finally eluted. After normalization, all DNA was pooled together and a purification step using magnetized beads in two steps to selectively remove fragments of DNA that are both too large and too small for sequencing. Following purification, each plate was quantified via qPCR using Life Technologies QuantStudio 6 Flex Instrument (Life Technologies). Finally, sequencing was performed using an Illumina HiSeq 1500 instrument.</p> <p><strong>Ancillary peer-reviewed manuscripts:</strong><br> <em>Genotyping protocols</em><br> Campbell NR, Harmon SA, Narum SR. 2015. Genotyping-in-Thousands by sequencing (GT-seq): A cost effective SNP genotyping method based on custom amplicon sequencing. Mol Ecol Resour, 15: 855-867. https://doi.org/10.1111/1755-0998.12357<br> <em>SNP loci reference</em><br> Collins EE, Hargrove JS, Delomas TA, Narum SR. 2020. Distribution of genetic variation underlying adult migration timing in steelhead of the Columbia River basin. Ecology and Evolution, 10(17): 9486-9502. https://doi.org/10.1002/ece3.6641&nbsp;&nbsp;</p> <p><strong>Data Use</strong>:<br> <em>License</em>: <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY 4.0</a>&nbsp; &nbsp;<br> <em>Recommended Citation</em>: Wooding AP, Narum SR, Pradhan DS. 2022. Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams (0.1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7055582</p> <p>Funding for this project is provided by&nbsp;US National Science Foundation and Idaho EPSCoR&nbsp;through award: OIA-1757324&nbsp;&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue for single nucleotide polymorphisms discovery (data set)

<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of <em>Festuca pratensis</em> and ten samples (corresponding to six genotypes) of <em>Lolium multiflorum</em> and fourteen samples of<em> Lolium perenne</em> (corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant <em>F. pratensis</em> transcripts were assembled by combining the contigs of all six genotypes based on orthology with the <em>Brachypodium distachyon </em>proteome. Similarly, <em>19,036</em> non-redundant<em> L. multiflorum</em> transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em> L. perenne</em> transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>

opencc-zeroFeb 2016View details →
zenodo40/100

F I G U R E 3 A in A low-density single nucleotide polymorphism panel for brown trout (Salmo trutta L.) suitable for exploring genetic diversity at a range of spatial scales

F I G U R E 3 A priori discriminant analysis of principal components (DAPC) plot of Camel trout. Each point represents the genotype of an individual fish, with centroids for each site labelled. Discriminant function 1 (DF1) is represented by the x axis, and discriminant function 2 (DF2) by the y-axis

opencc-by-4.0Nov 2022View details →
zenodo40/100

F I G U R E 1 in A low-density single nucleotide polymorphism panel for brown trout (Salmo trutta L.) suitable for exploring genetic diversity at a range of spatial scales

F I G U R E 1 Map showing the location of rivers sampled for brown trout within the UK, France and Ireland. The left panel shows the rivers used to assess the performance of the single nucleotide polymorphisms (SNP) panel at characterising genetic parameters within and outside the target region. The top right (blue) panel shows the locations of the four sampled rivers in Mount's Bay, Cornwall (Case Study 1). The bottom right (red) panel shows the location of the sample locations in the Camel catchment (Case Study 2). The red box within the bottom right panel gives the position of the impassable De Lank quarry site

opencc-by-4.0Nov 2022View details →
zenodo40/100

F I G U R E 2 A in A low-density single nucleotide polymorphism panel for brown trout (Salmo trutta L.) suitable for exploring genetic diversity at a range of spatial scales

F I G U R E 2 A priori discriminant analysis of principal components (DAPC) of trout genotypes from rivers flowing into Mount's Bay, Cornwall. Individuals are represented by individual points, with centroids for each river labelled. Discriminant function 1 (DF1) is represented by the x axis, and discriminant function 2 (DF2) by the y-axis

opencc-by-4.0Nov 2022View details →
zenodo40/100

F I G U R E 4 in A low-density single nucleotide polymorphism panel for brown trout (Salmo trutta L.) suitable for exploring genetic diversity at a range of spatial scales

F I G U R E 4 Correlation between geographic distance (km) against genetic distance (linear FST) for the trout samples from the River Camel. The red points represent those between the De Lank and all other sites, the black points for all pair-wise comparisons excluding the De Lank. Linear regression for all sites including the De Lank is given by the red line (r2 = 0.321, P = 0.231), and linear regression for all pair-wise sites excluding the De Lank is given by the black line (r2 = 0.658, P = 0.0671)

opencc-by-4.0Nov 2022View details →
dryad40/100

Single nucleotide polymorphism (SNPs) data for Scurria scurra, Scurria variabilis, Scurria ceciliana and Scurria araucana

<p>The distribution of genetic diversity is often heterogeneous in space, and it usually correlates with environmental transitions or historical processes that affect demography. The coast of Chile encompasses two biogeographic provinces and spans a broad environmental gradient together with oceanographic processes linked to coastal topography that can affect species' genetic diversity. Here, we evaluated the genetic connectivity and historical demography of four <em>Scurria</em> limpets, <em>S. scurra, S. variabilis, S. ceciliana</em> and <em>S. araucana</em>, between ca. 19° S and 53° S in the Chilean coast using genome-wide SNPs markers. Genetic structure varied among species which was evidenced by species-specific breaks together with two shared breaks. One of the shared breaks was located at 22–25° S and was observed in <em>S. araucana</em> and <em>S. variabilis</em>, while the second break around 31–34° S was shared by three <em>Scurria</em> species. Interestingly, the identified genetic breaks are also shared with other low-disperser invertebrates. Demographic histories show bottlenecks in <em>S. scurra</em> and <em>S. araucana</em> populations and recent population expansion in all species. The shared genetic breaks can be linked to oceanographic features acting as soft barriers to dispersal and also to historical climate, evidencing the utility of comparing multiple and sympatric species to understand the influence of a particular seascape on genetic diversity.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Supplementary material for: "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"

<p>Supplementary material for the publication &quot;Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations&quot;</p> <p>The realated preprint can be found at Research Square (<a href="https://doi.org/10.21203/rs.3.rs-861830/v1">https://doi.org/10.21203/rs.3.rs-861830/v1</a>)</p> <p>Supplementary file 1: Supplementary results, tables and figures.</p> <p>Supplementary file 2: MultiQC report.</p> <p>Supplementary file 3: Observer concordance of the visual filtering step.</p> <p>Supplementary file 4: Snakemake workflow and scripts.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Single-nucleotide polymorphisms in isolates Pyricularia oryzae isolates from rice and other hosts

<p>We investigated the presence of inter-lineage hybrids in a set of 886 isolates of the rice-infecting lineage of the rice blast fungus (<em>Pyricularia oryzae</em>).</p> <p>More details in https://doi.org/10.1101/2020.06.02.129296</p> <p>List of files:</p> <p>concatenation_5190SCOs_179isolates.fasta.gz: concatenated sequences of 5190 single copy orthologs from 179 isolates (used in Figures A and B, and computations of polymorphism and divergence in Appendix 1).</p> <p>concatenation_3686SNPs_415isolates.fasta.gz: concatenated polymorphisms identified in 415 isolates at 3686 genomic positions (used in Figure C, Appendix 1).</p> <p>concatenation_503889SNPs_96isolates.fasta.gz: concatenated polymorphisms identified in 96 isolates at 503889 genomic positions (used in Figures D and E, Appendix 1).</p> <p>orthology_table_with_70-15_genemodels.txt: Orthology table, with one orthogroup per line, one species per column, 16445 orthogroups, 179 isolates. The MGG column represents gene models from the 70-15 reference genome MG8, determined subsequently to orthology analysis by blastx of 70-15 sequences against isolate VT0027, which was the closest isolate in a neighbor-net network.</p>

opencc-by-4.0Sep 2021View details →
dryad40/100

Single nucleotide polymorphism (SNPs) data for Scurria scurra, Scurria variabilis, Scurria ceciliana and Scurria araucana

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad40/100

Assessing kinship detection: Single nucleotide polymorphism array density and estimator comparison in white-tailed deer

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad40/100

Data from: FSHB transcription is regulated by a novel 5’ distal enhancer containing a fertility-associated single nucleotide polymorphism

Open the record for dataset details and reuse information.

publicOct 2020View details →
zenodo36/100

A genomic data set of single‐nucleotide polymorphisms (SNPs) generated by ddRAD tag sequencing in Q. petraea (Matt.) Liebl. populations from Central-Eastern Europe and Balkan Peninsula

<p>This genomic dataset provides highly variable single-nucleotide polymorphism&nbsp;(SNP) markers from georeferenced natural <em>Quercus petraea</em> (Matt.) Liebl. populations collected in Bulgaria, Hungary, Romania, Serbia, Bosnia and Herzegovina, Kosovo and Albania. These SNP loci can be used to assess genetic diversity, differentiation, population structure, and can also be used to detect signatures of selection and local adaptation.</p>

opencc-by-4.0Jun 2020View details →
dryad36/100

Data from: Discovery and characterization of single nucleotide polymorphisms in Chinook salmon, Oncorhynchus tshawytscha

Molecular population genetics of non-model organisms has been dominated by the use of microsatellite loci over the last two decades. The availability of extensive genomic resources for many species is contributing to a transition to the use of single nucleotide polymorphisms (SNPs) for the study of many natural populations. Here we describe the discovery of a large number of SNPs in Chinook salmon, one of the world's most important fishery species, through large-scale Sanger sequencing of expressed sequence tag (EST) regions. More than 3MB of sequence was collected in a survey of variation in more than 131KB of unique genic regions, from more than 225 separate ESTs, in a diverse ascertainment panel of 24 salmon. This survey yielded 117 TaqMan (5' nuclease) assays, almost all from separate EST regions, which were validated in population samples from 5 major stocks of salmon from the three largest basins on the Pacific coast of the coterminous United States: the Sacramento, Klamath and Columbia Rivers. The proportion of these loci that was variable in each of these stocks ranged from 86.3 to 90.6% and the mean minor allele frequency ranged from 0.194 to 0.236. There was substantial differentiation between populations with these markers, with a mean FST estimate of 0.107, and values for individual loci ranging from 0 to 0.592. This substantial polymorphism and population-specific differentiation indicates that these markers will be broadly useful, including for both pedigree reconstruction and genetic stock identification applications.

opencc-zeroDec 2009View details →
dryad36/100

Data from: Discovery and characterization of single nucleotide polymorphisms in two anadromous alosine fishes of conservation concern

Freshwater habitat alteration and marine fisheries can affect anadromous fish species, and populations fluctuating in size elicit conservation concern and coordinated management. We describe the development and characterization of two sets of 96 single nucleotide polymorphism (SNP) assays for two species of anadromous alosine fishes, alewife and blueback herring (collectively known as river herring), that are native to the Atlantic coast of North America. We used data from high-throughput DNA sequencing to discover SNPs and then developed molecular genetic assays for genotyping sets of 96 individual loci in each species. The two sets of assays were validated with multiple populations that encompass both the geographic range and the known regional genetic stocks of both species. The SNP panels developed herein accurately resolved the genetic stock structure for alewife and blueback herring that was previously identified using microsatellites and assigned individuals to regional stock of origin with high accuracy. These genetic markers, which generate data that are easily shared and combined, will greatly facilitate ongoing conservation and management of river herring including genetic assignment of marine caught individuals to stock of origin.

opencc-zeroDec 2016View details →
dryad36/100

Data from: Evaluation of a single nucleotide polymorphism baseline for genetic stock identification of Chinook Salmon (Oncorhynchus tshawytscha) in the California Current Large Marine Ecosystem

Chinook Salmon is an economically and ecologically important species, and populations from the west coast of North America are a major component of fisheries in the North Pacific Ocean. The anadromous life history strategy of this species generates populations (or stocks) that typically are differentiated from neighboring populations. In many cases, it is desirable to discern the stock of origin of an individual fish or the stock composition of a mixed sample to monitor the stock-specific effects of anthropogenic impacts and alter management strategies accordingly. Genetic stock identification (GSI) provides such discrimination, and we describe here a novel GSI baseline composed of genotypes from more than 8000 individual fish from 69 distinct populations at 96 single nucleotide polymorphism (SNP) loci. The populations included in this baseline represent the likely sources for more than 99% of the salmon encountered in ocean fisheries of California and Oregon. This new genetic baseline permits GSI with the use of rapid and cost-effective SNP genotyping, and power analyses indicate that it provides very accurate identification of important stocks of Chinook Salmon. In an ocean fishery sample, GSI assignments of more than 1000 fish, with our baseline, were highly concordant (98.95%) at the reporting unit level with information from the physical tags recovered from the same fish. This SNP baseline represents an important advance in the technologies available to managers and researchers of this species.

opencc-zeroDec 2013View details →
zenodo36/100

Drosophila simulans VCF: The set of single nucleotide polymorphisms and insertion/deletions in a population of 170 Drosophila simulans lines.

<p>Heritable phenotypic variation in natural populations exceeds the levels predicted under mutation-selection balance where purifying selection removes variation. Balancing selection, inefficient or weak selection, polygenic adaptation, and non-equilibrium populations are all possible explanations for excess variation. Yet, available genomic data indicate an abundance of directional selection. One potential explanation is that fleeting directional selection drives beneficial mutations to high frequency in rapid waves resulting in many intermediate frequency haplotypes. This hypothesis is supported by the genomic data from a panel of 170 D. simulans genotypes established from a single stable population which show evidence for an abundance of incomplete soft sweeps. Demography, admixture, and balancing selection cannot entirely explain the patterns in these data, while transient selective sweeps can account for all the patterns of variation observed in this population. One interpretation is that constant environmental shifts rapidly change the optimal phenotype within Drosophila populations, leaving a signature of adaptive responses.</p>

opencc-by-4.0Sep 2016View details →
dryad36/100

Data from: Distances and their visualization in studies of spatial-temporal genetic variation using single nucleotide polymorphisms (SNPs)

<p>Distance measures are widely used for examining genetic structure in datasets that comprise many individuals scored for a very large number of attributes. Genotype datasets composed of single nucleotide polymorphisms (SNPs) typically contain bi-allelic scores for tens of thousands if not hundreds of thousands of loci.</p> <p>We examine the application of distance measures to SNP genotypes and sequence tag presence-absences (SilicoDArT) and use real datasets and simulated data to illustrate pitfalls in the application of genetic distances and their visualization.</p> <p>The datasets used to illustrate points in the associated review are provided here together with the R script used to analyse the data. Data are either simulated internal to this script or are SNP data generated as part of other studies and included as compressed binary files readily accessable by reading into R using R base function readRDS(). Refer to the analysis script for examples.</p>

opencc-zeroJan 2024View details →
dryad36/100

Genome-wide single nucleotide polymorphisms reveal the genetic diversity and population structure of Creole goats from northern Peru

<p>Goat farming constitutes a significant source of income for farmers in northern Peru. There is currently an absence of information about the genetics of Peruvian Creole goats that would enable us to understand their origins and genetic spread. The objective of this study was to estimate the genetic diversity of Creole goats from northern Peru using SNP markers. This study involved the collection of 192 male Creole goats from three key goat production regions in northern Peru. These goat samples were genotyped using the GGPGoat70k SNP panel. To explore the genetic influence of other breeds on Peruvian Creole goats, our dataset was combined with previously published SNP genotypes. External data set includes multiple breeds genotypes sampled from Argentina, Brazil, Spain, and Alpine breed from Italy, France, and Switzerland. After quality control 52,832 autosomal SNPs were used to assess genetic diversity in the Peruvian goats. For the population structure analysis of the merged data 20,513 common SNPs were used. Estimations for expected heterozygosity (H<sub>e</sub>), observed heterozygosity (H<sub>o</sub>), and inbreeding coefficient (F<sub>IS</sub>) were computed for the Peruvian groups. AMOVA, principal component analysis and ADMIXTURE were conducted to evaluate the population structure in the two data sets, Peru and merged. The results revealed a considerable genetic diversity, with H<sub>o</sub> values ranging from 0.40 to 0.41 for the Peruvian sampling groups, and inbreeding coefficient was notably low for Peruvian goat. The population structure analysis demonstrated a distinction (p&lt; 0.05) from other breeds. These findings suggest a level of genetic differentiation of the Peruvian goat population among other breeds, although further research is needed considering samples from other Peruvian areas. We expect this study will contribute to define genetic management strategies to prevent the loss of genetic diversity in Peruvian goat populations and for upcoming advancements in this field.</p>

opencc-zeroMay 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record