Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
294
datasets available to search
ShareScore release 0.7.1
Dataset results
294 results for “SNPs”
Linkage-independent SNPs in the Drosophila melanogaster Sussex LHM sample
<p>Unix code for running Plink program for generating a list of SNPs (single-nucleotide polymorphisms) which are independent of linkage diseqiulibrium. Used for later statistical analyses incorporating the number of independent tests made across the genome.</p>
2.8M SNPs Chinese Spring RefSeq v2.1 dataset
<p>VCF file of 2,799,166 single nucleotide polymorphism (SNP) markers positioned onto the Chinese Spring reference assembly RefSeq v2.1 developed by the International Wheat Genome Sequence Consortium (IWGSC; Zhu et al., 2021). These SNPs were lifted from the 1,000 wheat exome project, originally positioned onto RefSeq v1.0 (He et al., 2019). The SNP projection from RefSeq v1.0 onto RefSeq v2.1 was accomplished using LiftOff (Shumate and Salzberg, 2021).</p> <p>References</p> <p>He F, Pasam R, Shi F, Kant S, Keeble-Gagnere G, Kay P, Forrest K, Fritz A, Hucl P, Wiebe K, et al: <strong>Exome sequencing highlights the role of wild-relative introgression in shaping the adaptive landscape of the wheat genome.</strong> <em>Nature Genetics </em>2019, <strong>51:</strong>896-904.</p> <p>Shumate A, Salzberg SL: <strong>Liftoff: accurate mapping of gene annotations.</strong> <em>Bioinformatics </em>2021, <strong>37:</strong>1639-1643.</p> <p>Zhu T, Wang L, Rimbert H, Rodriguez JC, Deal KR, De Oliveira R, Choulet F, Keeble-Gagnère G, Tibbits J, Rogers J, et al: <strong>Optical maps refine the bread wheat Triticum aestivum cv. Chinese Spring genome assembly.</strong> <em>The Plant Journal </em>2021, <strong>107:</strong>303-314..</p>
Data from: Local adaptation (mostly) remains local: reassessing environmental associations of climate-related candidate SNPs in Arabidopsis halleri
<p>Numerous landscape genomic studies have identified single-nucleotide polymorphisms (SNPs) and genes potentially involved in local adaptation. Rarely, it has been explicitly evaluated whether these environmental associations also hold true beyond the populations studied. We tested whether putatively adaptive SNPs in <em>Arabidopsis</em> <em>halleri</em> (Brassicaceae), characterized in a previous study investigating local adaptation to a highly heterogeneous environment, show the same environmental associations in an independent, geographically enlarged set of 18 populations. We analysed new SNP data of 444 plants with the same methodology (partial Mantel tests, PMTs) as in the original study and additionally with a latent factor mixed model (LFMM) approach. Of the 74 candidate SNPs, 41% (PMTs) and 51% (LFMM) were associated with environmental factors in the independent data set. However, only 5% (PMTs) and 15% (LFMM) of the associations showed the same environment–allele relationships as in the original study. In total, we found 11 genes (31%) containing the same association in the original and independent data set. These can be considered prime candidate genes for environmental adaptation at a broader geographical scale. Our results suggest that selection pressures in highly heterogeneous alpine environments vary locally and signatures of selection are likely to be population-specific. Thus, genotype-by-environment interactions underlying adaptation are more heterogeneous and complex than is often assumed, which might represent a problem when testing for adaptation at specific loci.</p>
The list of SNPs
<p>This a long list = {{info for hb}, {info for Kr}, {info for gt}, {info for kni}};<br /> info for gene = list with elements {"chromosome name", global coordinate of polymorphic position, "nucleotide in the reference genome at this position", the list of nucleotides at this position for 216 lines, for example {"A","A"} (two letters for diploids) }</p> <p> </p>
Genome-wide estimation of linkage disequilibrium-independent SNPs in Drosophila melanogaster (Sussex LHM).
<p>Uses R to create SNP density across each chromosome arm. Uses Plink 1.9 to select independent SNPs with step sizes corresponding to chromosome density. Output data is combined_chromosomes_lhm_indep.txt a list of SNP IDs.<br> </p>
SNPs genotypes of southern beech Nothofagus dombeyi
<p>Geogenomics seeks to understand geological processes linked to lineage divergence. However, the mechanisms that conserve ancient signals despite gene flow are still unclear. In the southern beech, the deep lineage divergence produced by vicariant events is associated with ancient marine transgressions. We hereby evaluate the hypothesis that this divergence is maintained by diversifying selection. The lineage divergence using AMOVA, principal coordinate analysis, assignment tests, and multiple matrix regression analyses was assessed using chloroplast DNA and neutral and outlier SNPs. Several environmental variables were used to characterize potential within-species niche structuring and genotype-environment associations. Two deep-rooted latitudinally structured lineages resulted from cpDNA, the northern cluster being more genetically diverse than the southern one. Of the total of 2,943 SNPs, 33 were identified as outliers and produced two genetic clusters. Neutral SNPs yielded no structure by AMOVA, whereas higher (>75%) <em>F</em><sub>st</sub> values were obtained for cpDNA and outlier SNPs. Precipitation variables were mostly associated with population clusters and suggested two climatic niches, consisting of cold and dry in the south and more variable precipitation, temperature, and soil conditions in the north. Associations of genetic distance with environment and geography suggested IBD and IBE effects. Ancient lineage divergence in <em>N. dombeyi,</em> originally driven by vicariance, has been maintained by diversifying selection under distinct environmental conditions that also define distinct within-species niches. Deeply rooted phylogeographic breaks can be conserved in continuously distributed species in the absence of current geographic barriers. Yet physical gradients exert differential selective pressures, which are maintained in the face of potential gene flow. As a result, selection can lead to geographically localized and differentially adapted groups of populations that can be detected by a combination of traditional phylogeographic and novel genomic methods.</p>
Single nucleotide polymorphism (SNPs) data for Scurria scurra, Scurria variabilis, Scurria ceciliana and Scurria araucana
<p>The distribution of genetic diversity is often heterogeneous in space, and it usually correlates with environmental transitions or historical processes that affect demography. The coast of Chile encompasses two biogeographic provinces and spans a broad environmental gradient together with oceanographic processes linked to coastal topography that can affect species' genetic diversity. Here, we evaluated the genetic connectivity and historical demography of four <em>Scurria</em> limpets, <em>S. scurra, S. variabilis, S. ceciliana</em> and <em>S. araucana</em>, between ca. 19° S and 53° S in the Chilean coast using genome-wide SNPs markers. Genetic structure varied among species which was evidenced by species-specific breaks together with two shared breaks. One of the shared breaks was located at 22–25° S and was observed in <em>S. araucana</em> and <em>S. variabilis</em>, while the second break around 31–34° S was shared by three <em>Scurria</em> species. Interestingly, the identified genetic breaks are also shared with other low-disperser invertebrates. Demographic histories show bottlenecks in <em>S. scurra</em> and <em>S. araucana</em> populations and recent population expansion in all species. The shared genetic breaks can be linked to oceanographic features acting as soft barriers to dispersal and also to historical climate, evidencing the utility of comparing multiple and sympatric species to understand the influence of a particular seascape on genetic diversity.</p>
Eastern bettong (Bettongia gaimardi) reintroduced to Mulligan's Flat Woodland Sanctuary and Tidbinbilla Nature Reserve: DArT SNPs + individual information
<p>Incorporating genetic data into conservation programmes improves management outcomes, but the impact of different sample-grouping methods on genetic diversity analyses is poorly understood. To this end, the multi-source reintroduction of the eastern bettong (<em>Bettongia gaimardi</em>) was used as a long-term case study to investigate how sampling regimes may affect common genetic metrics, and hence management decisions. The dataset comprised 5307 SNPs sequenced across 263 individuals. Samples included 45 founders from five genetically distinct Tasmanian source regions, and 218 of their descendants captured during annual monitoring at Mulligan's Flat Woodland Sanctuary (MFWS; 121 samples across eight generations), and Tidbinbilla Nature Reserve (TNR; 97 samples across nine generations). The most management-informative sampling regime was found to be generational cohorts, providing detailed long-term trends in genetic diversity. When these generation-specific trends were not investigated, recent changes in population genetics were masked, and it became apparent that management recommendations would be less appropriate. The results also illuminated the importance of considering establishment and persistence as separate phases of a multi-source reintroduction. The establishment phase (useful for informing early adaptive management) should consist of no less than two generations, and continue until admixture is achieved (admixture defined here as >80% of individuals possessing >60% of source genotypes, with no one source composing >70% of >20% individuals' genotype) is achieved. This ensures that the persistence phase analyses of population trends remain minimally biased. Based on this case study, we recommend that emphasis be given to the value of generationally specific analyses, and that conservation programmes collect DNA samples throughout the establishment and persistence phases, and avoid collecting genetic samples only when analysis is imminent. We also recommend that population genetic analyses for multi-source reintroductions consider whether admixture has been achieved when calculating descriptive genetic metrics. </p>
Evidence for Correlations Between BMI-Associated SNPs and circRNAs
<p><strong>The datasets provided here are part of the study "Evidence for Correlations Between BMI-Associated SNPs and circRNAs" by Rajcsanyi et al.</strong></p> <p><strong>Abstract of the study:</strong></p> <p>Circular RNAs (circRNAs) are regulators of processes like adipogenesis. Their expression can be modulated by SNPs. We analysed links between BMI-associated SNPs and circRNAs. First, we detected an enrichment of BMI-associated SNPs on circRNA genomic loci in comparison to non-significant variants. Analysis of sex-stratified GWAS data revealed that circRNA genomic loci encompassed more genome-wide significant BMI-SNPs in females than in males. To explore if the enrichment is restricted to BMI, we investigated nine additional GWAS studies. We showed an enrichment of trait-associated SNPs in circRNAs for four analysed phenotypes (body height, chronic kidney disease, anorexia nervosa and autism spectrum disorder). To analyse the influence of BMI-affecting SNPs on circRNA levels in vitro, we examined rs4752856 located on hsa_circ_0022025. The analysis of heterozygous individuals revealed an increased level of circRNA derived from the BMI-increasing SNP allele. We conclude that genetic variation may affect the BMI partly through circRNAs.</p> <p><strong>Information regarding the datasets:</strong></p> <p>The data provided represents the analysed as well as generated data throughout the study. For further information about the datasets used, processed and generated, please see the study by Rajcsanyi et al.</p> <p><strong>circRNA datasets:</strong></p> <p>The analysed circRNA datasets were extracted from four circRNA databases (circAtlas v2.0, circBase, CIRCpediaV2 and circVAR) and were further processed to exclude internal duplicates and circRNAs derived from sex chromosomes. These original datasets have been downloaded from the following websites:</p> <p><em>circAtlas v2.0: http://159.226.67.237:8080/new/links.php</em></p> <p><em>circBase: http://www.circbase.org/cgi-bin/downloads.cgi</em></p> <p><em>CIRCpediaV2: http://yang-laboratory.com/circpedia/download</em></p> <p><em>circVAR: http://soft.bioinfo-minzhao.org/circvar/</em></p> <p><strong>GWAS datasets:</strong></p> <p>The original genome-wide association study (GWAS) summary statistics dataset of the BMI GWAS by Yengo et al. (2018) were classified into significant (P < 5*10<sup>-8</sup>) and non-significant (P >= 5*10<sup>-8</sup>) SNPs. A subsequent sensitivty analysis adjusted the P-value threshold of the non-significant SNPs to either P >= 5*10<sup>-5</sup>, P >= 5*10<sup>-6</sup> or 5*10<sup>-7</sup>.<br> Please note that these datasets are not provided in this repository, as these were solely classified and divided based on the SNPs' P-value. The same applied for all additional GWAS data solely divided based on the P-value (GWAS data for Anorexia nervosa, Autism spectrum disorder, etc.). The original and complete summary statistcs data can be obtained in the stated references below for each GWAS. Yet, the datasets generated for an approximation of the linkage disequilibrium based on the GWAS data by Yengo et al. (2018, BMI) are provided in this repository.<br> <br> <em>BMI and body height: Yengo et al. (2018)</em></p> <p><em>BMI sex-stratified: Pulit et al. (2019)</em></p> <p><em>Anorexia nervosa: Watson et al. (2018)</em></p> <p><em>Amyotropic lateral scerlosis: Iacoangeli et al. (2020)</em></p> <p><em>Autism spectrum disorder: Grove et al. (2019)</em></p> <p><em>Chronic kidney disease: Wuttke et al. 2019</em></p> <p><em>Epilepsy: ILAE consortium et al. 2018</em></p> <p><em>Heart Failure: Shah et al. 2020</em></p> <p><em>Pernicious anemia: Glanville et al. 2021</em></p> <p><em>Ulcerative Colitis: de Lange et al. 2017</em></p> <p><strong>Generated data:</strong></p> <p>The unprocessed output data of the study produced by the custom R script is provided here. Please note that the amount of information (rsID, P-value, Beta-value, allele frequency, circRNA_ID, circRNA strand, etc.) can vary between the output files due to differences in the data included in each circRNA and GWAS dataset. These dataset represent the raw and thus unprocessed output data. The results/counts presented in the study were obtained by further processing these output files. Further, the data produced by the SNaPshot assay are provided as well.</p>
Dataset: Synopsys, Inc. (SNPS) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
NanoValid D.5.47 Annex 1: Inter-laboratory comparison on measurand particle size of ~15 nm Lys-SNPs-B1 Silica Nanoparticles Particles
<p>An inter-laboratory comparison on the particle size, expressed as mean diameter <em>d</em>, of nanoscaled SiO<sub>2</sub> (#14 BAM Silica, ~15 nm diameter (see D.5.41/5.42)) has been performed. The majority of participants used Dynamic Light Scattering (DLS). A few used Electron Microscopy as method (SEM, TEM, T-SEM). Following methods had been applied by only one participant: Small Angle X-ray Scattering, Analytical Ultracentrifugation, Atomic Force Microscopy and Atomizer with electric mobility spectrometer (SMPS).</p> <p>The Task 5.4 of NanoValid is designed to test, compare and validate current methods to measure and characterize physicochemical properties of selected engineered nanoparticles. This will be achieved by inter-laboratory comparisons. The measurand of one of these round robins is <em>Particle size/Particle size distribution.</em> The measurements are to be accompanied by estimates of the uncertainties at a confidence level of 95%, deduced from the standard uncertainties. Therefore an uncertainty budget comprising statistical (Type A) and systematic (Type B) errors has to be established and delivered for each measurand. The inter-laboratry comparison protocol comprises two Annexes addressing the establishment of uncertainty budgets following GUM. The final goal of the comparison is to identify those methods of measurement which have potential as reference methods in pc characterization of nanoparticles for the determination of a given measurand (Task 5.4 of the NanoValid Project).</p>
Pooled DNA sequencing to identify SNPs associated with a major QTL for bacterial wilt resistance in Italian ryegrass (Lolium multiflorum Lam.)
<p>We used pooled DNA sequencing to characterize a major QTL for bacterial wilt resistance of Italian ryegrass and to develop inexpensive sequence-based markers to efficiently target resistance alleles for marker-assisted recurrent selection. From the mapping population segregating for the QTL, DNA of 44 of the most resistant and 44 of the most susceptible F<sub>1</sub> individuals were pooled and sequenced using the Illumina HiSeq2000 platform. Allele frequencies of 18 x 10<sup>6</sup> single nucleotide polymorphisms (SNP) were determined in the resistant and susceptible pool. A total of 271 SNPs on 140 scaffold sequences of the reference parental genome showed significantly different allele frequencies in both pools. We converted 44 selected SNPs to KASP markers, genetically mapped these proximal to the major QTL and thus validated their association with bacterial wilt resistance.</p>
43 longevity-associated SNPs genotyped in a Croatian sample of oldest-old individuals
<p>This dataset presents genotype data for 43 single nucleotide polymorphisms (SNPs) that have been genotyped in an anonymised sample of 314 oldest-old individuals (85+ years) from Croatia. The SNPs are located in or near candidate genes for longevity, and were selected from publicly available literature databases (PubMed and repositories specialized for human longevity such as https://genomics.senescence.info/longevity/, http://ageing-map.org/). They were selected based on their strong or repeatedly reported association with human longevity and involvement in various metabolic pathways. Genotyping was performed by Kompetitive Allele Specific PCR (KASP) on genomic DNA isolated from peripheral blood using the salting-out method. The dataset also contains recoding of the genotypes for each participant according to their association with longevity: a value of 2 was assigned to the homozygous genotype of longevity allele, a value of 1 to the heterozygous genotype, and a value of 0 to the homozygous genotype of an allele not associated with longevity in our sample. In cases where there were less than 10 of either homozygous genotypes, and in cases where a dominant or recessive coding gave a more significant result in further analyses, they were additionally recoded as binary variables with only the values 0 and 1, with heterozygote being added to the less common homozygote. This data was used to perform logistic regression analyses to create the best models for predicting survival to the ages of 90 and 95. Those models were then used to create genetic risk scores (here named genetic longevity scores, GLS) for predicting that phenotype, which are shown in this dataset as well. Information about the selected SNPs is also presented: rs code, nearest gene, chromosome position, and references for literature sources where association with longevity is reported; along with data that refers to the studied Croatian population: alleles (major/minor), minor allele frequencies (MAF), genotyping success rate, and HWE p-values. </p>
50000 SNPs for genomic prediction of ash dieback susceptibility in European Ash
<p>This file is <a href="https://www.nature.com/articles/s41559-019-1036-6">Stocks et al (2019)</a> Supplementary Table 7j with major allele (MAA) and minor allele (MIA) identities added.<br> Estimated effect sizes (EES) from genomic prediction model trained on the pool-seq data using the top 50000 SNPs from the pool-seq GWAS<br> Contig = Contig in BATG0.5 assembly <br> Pos = SNP location in contig <br> EES.MIA = Estimated effect size of minor allele <br> EES.MIA.SE = Standard Error of Estimated effect size of minor allele<br> EES.MAA = Estimated effect size of major allele<br> EES.MAA.SE = Standard Error of Estimated effect size of major allele <br> MIA = identity of minor allele <br> MAA = identity of major allele</p>
Data from: Assessment of coyote-wolf-dog admixture using ancestry-informative diagnostic SNPs
Open the record for dataset details and reuse information.
SNPs genotypes of southern beech Nothofagus dombeyi
Open the record for dataset details and reuse information.
Eastern bettong (Bettongia gaimardi) reintroduced to Mulligan's Flat Woodland Sanctuary and Tidbinbilla Nature Reserve: DArT SNPs + individual information
Open the record for dataset details and reuse information.
Data from: Local adaptation (mostly) remains local: reassessing environmental associations of climate-related candidate SNPs in Arabidopsis halleri
Open the record for dataset details and reuse information.
Single nucleotide polymorphism (SNPs) data for Scurria scurra, Scurria variabilis, Scurria ceciliana and Scurria araucana
Open the record for dataset details and reuse information.
SNPs derived from a common garden experiment across the biogeographic range of <em>Kelletia kelletii</em>
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.