Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

704

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

704 results for “nucleotides”

Learn how ShareScore rates datasets ↗
zenodo36/100

Single Nucleotide Polymorphisms (SNPs) identified from the whole genome sequences of hilsa shad (Tenualosa ilisha) of the Bay of Bengal

<p>The data file contains 792,939 isolated SNPs identified by discoSnp++ v2.3.x (Uricaru et al., 2015) from the whole genome sequence of T. ilisha of the Bay of Bengal. The central sequence of length 2k-1 is seen in upper case, while the flanking sequences are seen in lower case. SNP_higher/lower: one of the two alleles. id: id of the SNP (each SNP has a unique id).</p> <p>FOR SNPs:</p> <p>P_i:pos_Alt1/Alt2: Information about a ith SNP (If more than a unique SNP is found, the following format is used: P_1:pos_Alt1/Alt2,P_2:pos_Alt1/Alt2,...</p> <p>pos: position of the SNP with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>Alt1: One of the two alleles</p> <p>Alt2: the other</p> <p>FOR INDELs:</p> <p>P_1:pos_size_repeatSize</p> <p>pos: predicted position of the indel with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>size: predicted size of the indel</p> <p>repeatSize: Size of the longest sequence both prefix of the indel and prefix of the sequence located just after the insertion.</p> <p>high/low: sequence complexity. If the sequence if of low complexity (e.g. ATATATATATATATAT) this variable would be low</p> <p>nb_pol: number of polymorphism.</p> <p>left_unitig_length: size of the full left extension.</p> <p>right_unitig_length: size of the right extension.</p> <p>left_contig_length: size of the full left extension.</p> <p>right_contig_length: size of the right extension.</p> <p>C1: number of reads mapping the central upper case sequence from the first read set.</p> <p>C2: number of reads mapping the central upper case sequence from the second read set.</p> <p>Q1 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the first read set.</p> <p>Q2 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the second read set.</p> <p>G1: Genotype of the variant in the first read set.</p> <p>G2: Genotype of the variant in the second read set.</p> <p>rank: ranks the predictions according to their read coverage in each condition favoring SNPs that are discriminant between conditions.</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

Towards a rapid sequencing-based molecular surveillance and mosaicism investigation of Toxoplasma gondii (nucleotide alignment dataset)

<p>This dataset includes the nucleotide alignment of eight Toxoplasma gondii genome loci (Sag1 / Chromossome VIII, Gra6 / Chromossome X, PK1 / Chromossome VI, Sag3 / Chromossome XII, L363 / Chromossome VIIb, CB21-4 / Chromossome III, M102 / Chromossome VIIa, Sag2&nbsp;/ Chromossome VIII). Each alignment includes sequences from T. gondii reference strains (retrieved from ToxoDB) as well as sequences from multiple clinical strains (obtained by Sanger /&nbsp;Next-generation sequencing) of the collection of the&nbsp;National Reference Laboratory of Parasitic and Fungal Infections, Department of Infectious Diseases, National Institute of Health Dr. Ricardo Jorge, Portugal.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

Crown morphology in Norway spruce (Picea abies [Karst.] L.) as adaptation to mountainous environments is associated with single nucleotide polymorphisms (SNPs) in genes regulating seasonal growth rhythm

Trees growing at high altitude or latitude have to be adapted, amongst others, to the lower temperatures, a shorter vegetation period, heavier snow load and frost desiccation. Association between molecular genetic markers and climatic variables may provide evidence for the genetic control of climatic adaptation. With increasing genomic resources, several genes with importance to climatic adaptation are identified over a wide range of tree species. Commonly, circadian clock genes are linked to the adaptation to lower temperatures and especially to a shortened vegetation period, as they are regulating metabolic and phenological processes in the day-night shift and seasonal change. Potentially adaptive "candidate" genes associated with latitudinal and elevational gradients were identified in several Picea spp. Before molecular markers became available to study climatic adaptation, phenotypic traits measured in natural populations and/or common garden studies were used to search for their association with climate variables. In Norway spruce, the crown architecture is the most noticeable trait associated with altitude and the related environment. The mountainous narrow-crowned morphotype is characterised by superior resistance to snow breakage in regions with heavy snow fall. In total, the crown shape was assessed in 765 individual trees from mountainous regions in the Thuringian Forest, the Ore Mountains (Saxony) and Harz Mountains (Lower-Saxony/Saxony-Anhalt), and they were genotyped at 44 single nucleotide polymorphisms (SNPs) in 24 adaptive trait related candidate genes. Six SNPs in three genes, APETALA 2-like 3 (AP2L3), GIGANTEA (GI), and mitochondrial transcription termination factor (mTERF) were associated with variation in crown shape. GI has previously been identified in angiosperms and gymnosperms to be associated with temperature and growth cessation. Our results showed that crown morphology in Norway spruce is associated with genetic markers which are putatively involved in the complex process of genetic adaptation to climatic conditions at high altitudes.

opencc-zeroSep 2019View details →
zenodo36/100

RAD-seq generated single nucleotide polymorphisms resolve patterns of genetic diversity and structure of the freshwater mussel Ptychobranchus fasciolaris in glaciated and unglaciated regions of North America

<p>Included are the initial unfiltered SNP output from the STACKS pipeline, and the final filtered SNP dataset in VCF format used to do analysis in the manuscript titled "<span>RAD-seq generated single nucleotide polymorphisms resolve patterns of genetic diversity and structure of the freshwater mussel <em>Ptychobranchus fasciolaris </em>in glaciated and unglaciated regions of North America" which was submitted to <em>Hydrobiologia </em>in September 2024.</span></p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Supplementary Data to: "A single nucleotide mutation in DUOX2 gene causes some of the panda's unique metabolic phenotypes"

<p>This file contains the data set associated with the manuscript entitled: &quot;A single nucleotide mutation in the dual-oxidase 2 (<em>DUOX2</em>) gene causes some of the panda&rsquo;s unique metabolic phenotypes&quot;, National Science Review, DOI:&nbsp;<a href="http://dx.doi.org/10.1093/nsr/nwab125">10.1093/nsr/nwab125</a></p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Accompanying data set for the manuscript "REverSe TRanscrIptase Chain Termination (RESTRICT) for Selective Measurement of Nucleotide Analogs Used in HIV Care and Prevention"

<p>This data set contains all experimental and theoretical data included in the manuscript &quot; REverSe TRanscrIptase Chain Termination (RESTRICT) for Selective Measurement of Nucleotide Analogs Used in HIV Care and Prevention&quot;, namely:</p> <p>RESTRICT_model: MATLAB script for completing calculations in the RESTRICT theoretical model.</p> <p>Figure 2:</p> <ul> <li>Raw data from theoretical model showing contributions of individual model components, Kaff = 0.3</li> <li>Normalized data from theoretical model showing contributions of individual model components, Kaff = 0.3</li> <li>Experimental NRTI Drug Screen 180 nt TTCA 500 nM dNTP</li> </ul> <p>Figure 3:</p> <ul> <li>Experiment-dNTP-Concentration-Screen</li> <li>Theory-dNTP-Concentration-Screen</li> <li>Experiment-Template-Length-Screen</li> <li>Theory-Template-Length-Screen</li> <li>Experiment-Sequence-Screen</li> <li>Theory-Sequence-Screen</li> </ul> <p>Figure 4:</p> <ul> <li>Experimental-NRTI-Drug-Screen-90nt-TCAA-only</li> <li>Theory-TCAA90-Kaff=0point2</li> <li>Experimental-NRTI-Drug-Screen-GGCA-only</li> <li>Theory-GGCA180-Kaff=0point2</li> </ul> <p>Figure 5:</p> <ul> <li>GGCA-vs-TTCA-Specificity-Analysis</li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Vascular KATP channel structural dynamics reveal regulatory mechanism by Mg-nucleotides

<p>MD simulation data for vascular KATP channel focusing on Kir6.1-pore and SUR2B in the presence and absence of MgADP.&nbsp;</p>

opencc-by-4.0Oct 2021View details →
dryad36/100

High-Throughput-Methyl-Reading (HTMR) assay: A solution based on nucleotide methyl-binding proteins enables large-scale screening for DNA/RNA methyltransferases and demethylases

<p>Epigenetic therapy has significant potential for cancer treatment. However, few small potent molecules have been identified against DNA or RNA modification regulatory proteins. Current approaches for activity detection of DNA/RNA methyltransferases and demethylases are time-consuming and labor-intensive, making it difficult to subject them to high-throughput screening. Here, we developed a fluorescence polarization-based "High-Throughput Methyl Reading" (HTMR) assay to implement large-scale compound screening for DNA/RNA methyltransferases and demethylases-DNMTs, TETs, ALKBH5, and METTL3/METTL14. This assay is simple to perform in a mix-and-read manner by adding the methyl-binding proteins MBD1 or YTHDF1. The proteins can be used to distinguish FAM-labelled substrates or product oligonucleotides with different methylation statuses catalyzed by enzymes. Therefore, the extent of the enzymatic reactions can be coupled with the variation of FP binding signals. Furthermore, this assay can be effectively used to conduct a cofactor competition study. Based on the assay, we identified two natural products as candidate compounds for DNMT1 and ALKBH5. In summary, this study outlines a powerful homogeneous approach for high-throughput screening and evaluating enzymatic activity for DNA/RNA methyltransferases and demethylases that is cheap, easy, quick, and highly sensitive.</p>

opencc-zeroOct 2021View details →
zenodo36/100

S and Z loci nucleotide sequences of Lolium multiflorum cultivar Rabiosa

<p>Nucleotide sequences of two scaffolds spanning the&nbsp;<em>S</em>-locus and two scaffolds the&nbsp;<em>Z</em>-locus in&nbsp;<em>Lolium multiflorum</em>&nbsp;cultivar Rabiosa</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

GRAND-SLAM analysis of simulated nucleotide conversion in Illumina TruSeq data sets for grandRescue

<p>These are processed data sets from the simulation of nucleotide conversions (T&gt;C) in single-end and paired-end Illumina TruSeq reads for the purpose of investigating 4sU-induced mapping impairment by read lengths and library preparation methods and the potential of grandRescue to alleviate these effects.</p> <p>The original data set is from: Sarantopoulou, D. <em>et al. </em>(https://doi.org/10.1038/s41598-019-49889-1)</p> <p>GEO Accession:GSE124167 (samples: GSM3523316 - GSM3523318)</p> <p>&nbsp;</p> <p>The zip files contain the full output from the processing pipeline (including the mapped reads, the scripts to run the pipeline and the output) for single-end (R1) and paired-end before and after rescue. The *.tsv.gz files are the GRAND-SLAM output tables.</p> <p><br> To generate the GRAND-SLAM output yourself, first prepare the mouse genomes. Then run the following command with the respective cit-files, prefixes (*.cit) and genome:</p> <p>gedi -e Slam -trim5p 15 -reads *.cit -genomic m.ens102 -prefix grandslam_t15/* -plot&nbsp; -D -modelall</p> <p>To generate the cit file you have to modify the first lines in start.bash to match the paths on your file system, and then run it.</p> <p>&nbsp;</p> <p>Software versions:</p> <p>&nbsp;&nbsp;&nbsp; gedi toolkit 1.0.5<br> &nbsp;&nbsp;&nbsp; GRAND-SLAM 2.0.7<br> &nbsp;&nbsp;&nbsp; STAR version 2.7.10b</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

GRAND-SLAM analysis of simulated nucleotide conversion in QuantSeq data sets for grandRescue

<p>These are processed data sets from the simulation of nucleotide conversions (T&gt;C) in QuantSeq reads for the purpose of investigating 4sU-induced mapping impairment by read lengths and library preparation methods and the potential of grandRescue to alleviate these effects.</p> <p>The original data set is from: Lee, J. W. <em>et al. </em>(https://doi.org/10.1038/s41586-019-1004-y)</p> <p>GEO Accession: GSE109480 (Samples: GSM2944116 &ndash; GSM2944120)</p> <p>&nbsp;</p> <p>The zip files contain the full output from the processing pipeline (including the mapped reads, the scripts to run the pipeline and the output) before and after rescue. The *.tsv.gz files are the GRAND-SLAM output tables.</p> <p><br> To generate the GRAND-SLAM output yourself, first prepare the mouse genome. Then run the following command with the respective cit-files, prefixes (*.cit) and genome:</p> <p>gedi -e Slam -trim5p 15 -reads *.cit -genomic m.ens102 -prefix grandslam_t15/* -plot&nbsp; -D -modelall</p> <p>To generate the cit file you have to modify the first lines in start.bash to match the paths on your file system, and then run it.</p> <p>&nbsp;</p> <p>Software versions:</p> <p>&nbsp;&nbsp;&nbsp; gedi toolkit 1.0.5<br> &nbsp;&nbsp;&nbsp; GRAND-SLAM 2.0.7<br> &nbsp;&nbsp;&nbsp; STAR version 2.7.10b</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Xenon's sedative effect is at least partly mediated by interaction with the cyclic nucleotide-binding domain (CNBD) of HCN2 channels expressed by thalamocortical neurons of the ventrobasal nucleus in mice

<p>Xenon (70%xenon, 30%O2) application in the last 10 min of the open field test manage to sedate wild-type mice but not HCN2EA. Wild-type mice in the xenon_wild-type video can be seen moving freely in the early minutes but the general activity starts to decrease sharply in the last 2 minutes and is absent between min 9 and 10 while HCN2EA mice in the xenon_HCN2EA can be seen moving throughout the entire time of the xenon gas mixture application.</p>

opencc-by-4.0Apr 2023View details →
dryad36/100

Data for: Dietary nucleotides can prevent glucocorticoid-induced telomere attrition in a fast-growing wild vertebrate

<p><span>T<span>elomeres are chromosome protectors that shorten during eukaryotic cell replication and in stressful conditions. Developing individuals are </span>susceptible <span>to telomere erosion when their growth is fast and resources are limited. This is critical because the rate of telomere attrition in early life is linked to health and life span of adults. The metabolic telomere attrition hypothesis (MeTA) suggests that telomere dynamics can respond to biochemical signals conveying information about the organism's energetic state. Among these signals are glucocorticoids, hormones that promote catabolic processes, potentially impairing costly telomere maintenance, and nucleotides, which activate anabolic pathways through the cellular enzyme target of rapamycin (TOR), thus preventing telomere attrition. During the energetically demanding growth phase, the regulation of telomeres in response to two contrasting signals—one promoting telomere maintenance and the other attrition—provides an ideal experimental setting to test the MeTa. We studied nestlings of a rapidly developing free-living passerine, the great tit (<em>Parus</em> <em>major</em>), that either received glucocorticoids (Cort-chicks), nucleotides (Nuc-chicks), or a combination of both (NucCort-chicks), comparing these with controls (Cnt-chicks). As expected, Cort-chicks showed telomere attrition, while NucCort- and Nuc-chicks did not. NucCort-chicks was the only group showing increased expression of a proxy for TOR activation (the gene telo2), of mitochondrial enzymes linked to ATP production (cytochrome oxidase and ATP-synthase) and a higher efficiency in aerobically producing ATP. NucCort-chicks had also a higher expression of telomere maintenance genes (shelterin protein TERF2 and telomerase TERT) and of enzymatic antioxidant genes (glutathione peroxidase and superoxide dismutase). The findings show </span>that nucleotide availability is crucial for preventing telomere erosion during fast growth in stressful environments. </span></p>

opencc-zeroAug 2023View details →
ClinicalTrials.gov36/100

Safety and Efficacy of Reduced Versus Standard Dose Efavirenz (EFV) Plus Two Nucleotides in Antiretroviral-naïve Adults.

ClinicalTrials.gov study NCT01011413. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad36/100

Nucleotide alignment and phylogenetic tree illustrating tick-derived Mycoplasma cynos

Open the record for dataset details and reuse information.

publicMay 2025View details →
dryad36/100

Data for: Dietary nucleotides can prevent glucocorticoid-induced telomere attrition in a fast-growing wild vertebrate

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad36/100

Data from: Distances and their visualization in studies of spatial-temporal genetic variation using single nucleotide polymorphisms (SNPs)

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad36/100

Data from: Evaluation of a single nucleotide polymorphism baseline for genetic stock identification of Chinook Salmon (Oncorhynchus tshawytscha) in the California Current Large Marine Ecosystem

Open the record for dataset details and reuse information.

publicMar 2015View details →
dryad36/100

Data from: Discovery and characterization of single nucleotide polymorphisms in coho salmon, Oncorhynchus kisutch

Open the record for dataset details and reuse information.

publicMay 2015View details →
dryad36/100

Single nucleotide polymorphism genotypes for the Australian blackspot shark and the milk shark in Northern Australian waters

Open the record for dataset details and reuse information.

publicNov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record