Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

704

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

704 results for “nucleotides”

Learn how ShareScore rates datasets ↗
zenodo32/100

SUPPLEMENTARY FIGURE 2. Tree generated from the nucleotide sequence for the mitochondrial gene region, igr1–cox1 in A taxonomic revision of Anthothela (Octocorallia: Scleraxonia: Anthothelidae) and related genera, with the addition of new taxa, using morphological and molecular data

SUPPLEMENTARY FIGURE 2. Tree generated from the nucleotide sequence for the mitochondrial gene region, igr1–cox1 of Anthothela-like specimens. Bayesian posterior probabilities shown above branch, ML bootstrap values below branch; HKY+G (Bayesian results split freq = 0.0019, 10000000 gen, burnin=25000). (* indicates nodes present only in Bayesian analysis).

opennotspecifiedDec 2017View details →
zenodo32/100

Data and scripts supporting "Metabolic profiling of patient-derived organoids reveals nucleotide synthesis as a metabolic vulnerability in malignant rhabdoid tumors"

<p>This submission contains the processed sequencing data and analysis scripts, as well as the raw and processed LC-MS metabolomics files accompanying our manuscript "Metabolic profiling of patient-derived organoids reveals nucleotide synthesis as a metabolic vulnerability in malignant rhabdoid tumors" (Cell Reports Medicine, 2024)</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Single-nucleotide-resolution genomic maps of O6-methylguanine from the glioblastoma drug temozolomide

<p>Temozolomide kills cancer cells by forming <em>O</em><sup>6</sup>-methylguanine (<em>O</em><sup>6</sup>-MeG), which leads to cell cycle arrest and apoptosis. However,<em> O</em><sup>6</sup>-MeG repair by <em>O</em><sup>6</sup>-methylguanine-DNA methyltransferase (MGMT) contributes to drug resistance. Characterizing genomic profiles of <em>O</em><sup>6</sup>-MeG could elucidate how <em>O</em><sup>6</sup>-MeG accumulation is influenced by repair, but there are no methods to map genomic locations of <em>O</em><sup>6</sup>-MeG. Here, we developed an immunoprecipitation- and polymerase-stalling-based method, termed <em>O</em><sup>6</sup>-MeG-seq, to locate <em>O</em><sup>6</sup>-MeG across the whole genome at single-nucleotide resolution. We analyzed <em>O</em><sup>6</sup>-MeG formation and repair with regards to sequence contexts and functional genomic regions in a glioblastoma-derived cell line and evaluated the impact of MGMT expression.&nbsp;<em>O</em><sup>6</sup>-MeG signatures were highly similar to mutational signatures from patients previously treated with temozolomide. Furthermore, MGMT did not preferentially repair <em>O</em><sup>6</sup>-MeG with respect to sequence context, chromatin state or gene expression level, however, may protect oncogenes from mutations. Finally, we found an MGMT-independent strand bias in <em>O</em><sup>6</sup>-MeG accumulation in highly expressed genes. These data provide high resolution insight on how <em>O</em><sup>6</sup>-MeG formation and repair are impacted by genome structure and nucleotide sequence. Further, <em>O</em><sup>6</sup>-MeG-seq is expected to enable future studies of DNA modification signatures as diagnostic markers for addressing drug resistance and preventing secondary cancers.</p>

opencc-by-4.0Oct 2024View details →
dryad32/100

Nucleotide sequences in Procambarus clarkii

<p>Crayfish is a model for studying the effect of light on locomotor activity and neuroendocrine functions. In this study, we have described 62 transcripts from the pleonal nerve cord of the crayfish, using bioinformatics tools that identify phylogenetic families of genes related to the light interaction in the extraretinal photoreceptors. We deposited all sequencing data in the GenBank database. Here shows supplemental data from the freshwater crayfish <i>Procambarus clarkii</i>. The results suggest that the genes related to ocular and extraocular light perception in the crayfish <i>P. clarkii</i> use common biosynthesis pathways and phototransduction cascades.</p>

opencc-zeroNov 2021View details →
dryad32/100

Nucleotide substitutions during speciation may explain substitution rate variation

<p><span>Although molecular mechanisms associated with the generation of mutations are highly conserved across taxa, there is widespread variation in mutation rates between evolutionary lineages. When phylogenies are reconstructed based on nucleotide sequences, such variation is typically accounted for by the assumption of a relaxed molecular clock, which is just a statistical distribution of mutation rates without much underlying biological mechanism. Here, we propose that variation in accumulated mutations may be partly explained by an elevated mutation rate during speciation. Using simulations, we show how shifting mutations from branches to speciation events impacts inference of branching times in phylogenetic reconstruction. Furthermore, the resulting nucleotide alignments are better described by a relaxed than by a strict molecular clock. Thus, elevated mutation rates during speciation potentially explain part of the variation in substitution rates that is observed across the tree of life. </span></p>

opencc-zeroNov 2021View details →
zenodo32/100

Western redcedar single nucleotide polymorphism (SNP) genotyping data for genomic selection and population genetics

<p>Western redcedar (<em>Thuja plicata</em>) Single Nucleotide Polymorphism (SNP) data in Variant Call Format (VCF) for genomic selection training and target populations, genomic selection parents, and self-fertilized (selfing) lines, comprising 4,833 trees.</p> <p>Targeted sequencing-based genotyping was done by Capture-Seq methodology at Rapid Genomics (Neves est al. 2013). A set of 57,000 probes as designed for initial marker discovery, from which a panel of 20,858 probes was selected for genotyping. A set of transcriptomes (Shalev et al. 2018) (PRJNA704616) was aligned to the reference genome to identify SNPs. Candidate probes (120 nt) were initially designed in silico and 57,000 selected by removing candidates with poor base composition for hybridization (GC content &lt;0.2 and &gt;0.6, high G content &gt;0.2 and long homopolymers &gt;7), followed by removing probes aligning to more than one position on the reference genome (&ge;90% identity and length). The 57,000 probes represent 14,517 scaffolds (average 3.9 probes/scaffold), with 37,275 targeting at least one SNP and 19,725 mapping to intergenic regions not containing pre-identified SNPs. A set of 128 individuals were selected to validate the 57,000 probe panel and associated polymorphisms. Genomic DNA (0.5 ug) was fragmented (mean size 300 bp), followed by repair of ends, phosphorylation, adenylation, ligation of Illumina compatible adapters containing 8bp indexes and 5&rsquo; T-overhang, and 10 cycles PCR amplification with universal primers to produce sequencing-ready libraries. Libraries were quantified using PicoGreen. Libraries from 16 samples were pooled, hybridized to the 120 nt RNA probes following Agilent&rsquo;s SureSelect Target Enrichment System (Agilent Technologies) and sequenced on an Illumina HiSeq X machine with paired-end 150bp cycle for an average sequencing depth per sample of 15X. Sequence data were aligned to the reference genome with BWA-MEM (http://arxiv.org/abs/1303.3997) and sets of four samples were combined to increase sequencing depth for identifying markers. Putative SNPs were identified using Freebayes (http://arxiv.org/abs/1207.3907) in 150bp on either side of the 57,000 probes and filtered probes that had more than 17 SNPs per 420 bp target region (150bp + 120bp + 150bp). The sequencing depth of the probes was used to select the final set of 20,885 probes, removing probes on both sides of the distribution (low and high sequencing depth), for Capture-Seq on the remainder of the samples.</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

FIGURE. Phylogenetic tree derived from Bayesian analysis, based on nrLSU data. Posterior probability (PP> 0.95) values from the Bayesian analysis are added at the nodes. The scale bar represents the number of nucleotide changes per site. (T) indicates the type specimen for this species. The new species are in bold. in Four new species of Entoloma (Entolomataceae, Agaricomycetes) subgenera Cyanula and Claudopus from Vietnam and their phylogenetic position

FIGURE. Phylogenetic tree derived from Bayesian analysis, based on nrLSU data. Posterior probability (PP&gt; 0.95) values from the Bayesian analysis are added at the nodes. The scale bar represents the number of nucleotide changes per site. (T) indicates the type specimen for this species. The new species are in bold.

opennotspecifiedJun 2022View details →
zenodo32/100

Fig. 2 in Inferring Ancestry and Divergence Events in a Forest Pest Using Low-Density Single-Nucleotide Polymorphisms

Fig. 2. Population tree inferred by TreeMix using all 42 sampling sites across the mountain pine beetle range. Terminal branches are labeled using locality information (see Supp Fig. 1 [online only] for further detail) and migration arrows are colored according to migration weight.

opennotspecifiedNov 2018View details →
zenodo32/100

Fig. 3. Scenario 9 in Inferring Ancestry and Divergence Events in a Forest Pest Using Low-Density Single-Nucleotide Polymorphisms

Fig. 3. Scenario 9 was identified as the 'best' from 11 competing phylogeographic scenarios. This scenario represents an east-to-west colonization route with stable population sizes in which cluster 4 is derived from clusters 1 and 3 through admixture. Terminal branch labels are consistent with the five STRUCTURE clusters identified (Fig. 1). T = time expressed as number of generations assuming one generation per year; R = inferred migration rate.

opennotspecifiedNov 2018View details →
zenodo32/100

Fig. 1 in Inferring Ancestry and Divergence Events in a Forest Pest Using Low-Density Single-Nucleotide Polymorphisms

Fig. 1. Map of mountain pine beetle sites throughout western North America. Each site is represented by a pie chart depicting the average assignment (Q values) of individuals to one of five STRUCTURE clusters (K = 5).

opennotspecifiedNov 2018View details →
dryad32/100

Unlocking the grain quality enigma: A KASP-driven voyage through bread wheat's quantitative trait nucleotides under heat adversity

<p>Heat stress is a critical factor affecting global wheat production and productivity. In this study, out of 500 studied accessions a diverse panel of 126 wheat genotypes grown under twelve distinct environmental conditions was analyzed. Using 35K single-nucleotide polymorphism (SNP) genotyping assays and trait data on five biochemical parameters, including grain protein content (GPC), grain amylose content (GAC), grain total soluble sugars (TSS), grain iron (Fe), and zinc (Zn) content, six multi-locus GWAS models were employed for association analysis. This revealed 67 significantly associated QTNs linked to grain quality parameters, explaining phenotypic variations ranging from 3% to 44% under heat stress conditions. By considering the results in consensus to at least three GWAS models and three locations, the final QTNs were reduced to 17, with 14 being novel findings. Notably, two novel markers, AX-94461119 (chromosome 6A) and AX-95220192 (chromosome 7D), associated with grain iron and zinc, respectively, were validated through KASP approach. Candidate genes, such as chaperonin Cpn60/GroEL/TCP-1 family, P-loop containing nucleoside triphosphate hydrolases (NTPases), Bowman-Birk type proteinase inhibitor (BBI), and NPSN13 protein, were identified from the associated genomic regions, which could be potentially targeted for improving quality traits and heat tolerance in wheat.</p>

opencc-zeroMay 2024View details →
zenodo32/100

Complete dataset used in the study "A statistical approach to coronavirus classification based on nucleotide distributions"

<p>This is a complete dataset of parameters used in coronavirus classification based on distributions of nucleotide sequences as described in the study "A statistical approach to coronavirus classification based on nucleotide distributions"</p>

opencc-by-4.0May 2024View details →
zenodo32/100

SASDM95 – Candida albicans Ras-like protein 1 in complex with the guanine nucleotide exchange factor region of cell division control protein 25

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

SASDM85 – GTP-binding domain of Candida albicans Ras-like protein 1 in complex with the guanine nucleotide exchange factor region of cell division control protein 25

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

SASDM75 – Ras-like protein 1 guanine nucleotide exchange factor region of Candida albicans cell division control protein 25

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
dryad32/100

Data from: Single nucleotide polymorphisms across a species' range: implications for conservation studies of Pacific salmon

Studies of the oceanic and near-shore distributions of Pacific salmon, whose migrations typically span thousands of kilometers, have become increasingly valuable in the presence of climate change, increasing hatchery production, and potentially high rates of bycatch in offshore fisheries. Genetics data offer considerable insights into both the migratory routes as well as the evolutionary histories of the species. However, these types of studies require extensive datasets from spawning populations originating from across the species? range. Single nucleotide polymorphisms (SNPs) have been particularly amenable for multi-national applications because they are easily shared, require little inter-laboratory standardization, and can be assayed through increasingly efficient technologies. Here we discuss the development of a dataset for 114 populations of chum salmon through a collaboration among North American and Asian researchers, termed PacSNP. PacSNP is focused on developing the database and applying it to problems of international interest. A dataset spanning the entire range of species provides a unique opportunity to examine patterns of variability, and we review issues associated with SNP development. We found evidence of ascertainment bias within the dataset, variable linkage relationships between SNPs associated with ancestral groupings, and outlier loci with alleles associated with latitude.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks

An Illumina Infinium SNP genotyping array was constructed for European white oaks. Six individuals of Quercus petraea and Q. robur were considered for SNP discovery using both previously obtained Sanger sequences across 676 gene regions (1371 in vitro SNPs) and Roche 454 technology sequences from 5112 contigs (6542 putative in silico SNPs). The 7913 SNPs were genotyped across the six parental individuals, full-sib progenies (one within each species and two interspecific crosses between Q. petraea and Q. robur) and three natural populations from south-western France that included two additional interfertile white oak species (Q. pubescens and Q. pyrenaica). The genotyping success rate in mapping populations was 80.4% overall and 72.4% for polymorphic SNPs. In natural populations, these figures were lower (54.8% and 51.9%, respectively). Illumina genotype clusters with compression (shift of clusters on the normalized x-axis) were detected in ~25% of the successfully genotyped SNPs and may be due to the presence of paralogues. Compressed clusters were significantly more frequent for SNPs showing a priori incorrect Illumina genotypes, suggesting that they should be considered with caution or discarded. Altogether, these results show a high experimental error rate for the Infinium array (between 15% and 20% of SNPs potentially unreliable and 10% when excluding all compressed clusters), and recommendations are proposed when applying this type of high-throughput technique. Finally, results on diversity levels and shared polymorphisms across targeted white oaks and more distant species of the Quercus genus are discussed, and perspectives for future comparative studies are proposed.

opencc-zeroDec 2014View details →
zenodo32/100

FIGURE 1 in High nucleotide divergence in a dimorphic parasite with disparate hosts

FIGURE 1. (A) Schematic line drawing of the secondary structure of the 5'-half of LSU 28S rRNA from the beetle Tenebrio sp. (redrawn from Gillespie et al. 2004). The shaded region shows the expansion segment D2 that was analyzed in this study. (B) Secondary structure model of the expansion segment D2 of the nuclear large subunit (28S) rRNA gene from Caenocholax fenyesi texensis (Texas locale). Bolded nucleotides depict variable characters in C. f. waloffi (Vera Cruz locale) with alternative states coloured red (and green where variation across individuals was detected). Blue helices depict structures strongly expanded in Strepsiptera. Helices in dark-lined boxes depict conserved helices across all arthropods. Helices in shaded boxes depict basepairs that are supported by both comparative evidence and thermodynamics. Basepairs within dash boxes illustrate support for one model (thermodynamics) that is not supported by comparative evidence across all sequences. Base-pairing is indicated as follows: standard canonical pairs by lines (C-G, G-C, A-U, U-A); wobble pairs by dots (G·U, U·G.); A-G pairs by open circles (A°G, G°A); other non-canonical pairs by filled circles (e.g. C•A). Gaps (-) represent insertions or deletions. Diagram was generated manually in Adobe Illustrator.

opennotspecifiedNov 2007View details →
dryad32/100

Improving Phylogenies Based on Average Nucleotide Identity, Incorporating Saturation Correction and Non-Parametric Bootstrap Support

<p>Whole genome comparisons based on Average Nucleotide Identities (ANI) and the Genome-to-genome distance calculator have risen to prominence in rapidly classifying prokaryotic taxa using whole genome sequences. Some implementations have even been proposed as a new standard in species classification and have become a common technique for papers describing newly sequenced genomes. However, attempts to apply whole genome divergence data to delineation of higher taxonomic units and to phylogenetic inference have had difficulty matching those produced by more complex phylogenetic methods. We present a novel method for generating statistically supported phylogenies of archaeal and bacterial groups using a combined ANI and alignment fraction-based metric. For the test cases to which we applied the developed approach we obtained results comparable with other methodologies up to at least the family-level.  The developed method uses non-parametric bootstrapping to gauge support for inferred groups.  This method offers the opportunity to make use of whole-genome comparison data, that are already being generated, to quickly produce phylogenies including support for inferred groups. Additionally, the developed ANI methodology can assist classification of higher taxonomic groups.<br> <br> Included herein are supplemental materials, and all whole genome datasets used throughout the construction of this work.</p>

opencc-zeroJul 2021View details →
dryad32/100

Analysis of RNA-seq, DNA target enrichment, and Sanger nucleotide sequence data resolves deep splits in the phylogeny of cuckoo wasps (Hymenoptera: Chrysididae)

<p>The wasp family Chrysididae (cuckoo wasps, gold wasps) comprises exclusively parasitoid and kleptoparasitic species, many of which feature a stunning iridescent coloration and phenotypic adaptations to their parasitic life style. Previous attempts to infer phylogenetic relationships among the family's major lineages (subfamilies, tribes, genera) based on Sanger sequence data were insufficient to statistically resolve the monophyly and the phylogenetic position of the subfamily Amiseginae and the phylogenetic relationships among the tribes Allocoeliini, Chrysidini, Elampini, and Parnopini (Chrysidinae). Here, we present a phylogeny inferred from nucleotide sequence data of 492 nuclear single-copy genes (230,915 aligned amino acid sites) from 94 species of Chrysidoidea (representing Bethylidae, Chrysididae, Dryinidae, Plumariidae) and 45 outgroup species by combining RNA-seq and DNA target enrichment data. We find support for Amiseginae being more closely related to Cleptinae than to Chrysidinae. Furthermore, we find strong support for Allocoeliini being the sister lineage of all remaining Chrysidinae, while Elampini represent the sister lineage of Chrysidini and Parnopini. Our study corroborates results from a recent phylogenomic investigation which revealed Chrysidoidea as likely paraphyletic</p>

opencc-zeroOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record