Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
96
datasets available to search
ShareScore release 0.9.0
Dataset results
96 results for “nucleotide sequences”
R notebooks to reproduce all analyses from the manuscript "grandR: a comprehensive package for nucleotide conversion sequencing data analysis"
<p>This package contains all R notebooks to reproduce the analyses from our manuscript "grandR: a comprehensive package for nucleotide conversion sequencing data analysis".</p> <p>In the zip file you find</p> <ul> <li>several rds files in the data folder: They contain grandR objects of both simulated and real SLAM-seq data sets. You can delete them and create them again by either just "knitting" the notebooks (which will generate all data necessary for this notebook and save it into the data folder), or by executing the generateAllDataFiles.R script ("Rscript generateAllDataFiles.R"), which will generate all rds files that do not exist).</li> <li>several R notebooks (Rmd): "Knitting" them will generate all figures from the manuscript. Without the data files (rds), this will be slow!</li> <li>knit_all.bash: Execute to "knit" all notebooks</li> <li>clean.bash: Clear the output of "knitting" the notebooks</li> </ul> <p> </p>
indicate branches. above MrBayes numbers by inferred The . supports Ixodes of probability subgenera 22 posterior the of Inference 16 from Bayesian ticks of indicate genomes mitochondrial branches below 40 numbers of The sequences. RAxML nucleotide by the inferred from support inferred bootstrap Phylogenies Likelihood . 2 FIGURE Maximum in A new subgenus, Australixodes n. subgen. (Acari: Ixodidae), for the kiwi tick, Ixodes anatis Chilton, 1904, and validation of the subgenus Coxixodes Schulze, 1941 with a phylogeny of 16 of the 22 subgenera of Ixodes Latreille, 1795 from entire mitochondrial genome sequences
indicate branches. above MrBayes numbers by inferred The . supports Ixodes of probability subgenera 22 posterior the of Inference 16 from Bayesian ticks of indicate genomes mitochondrial branches below 40 numbers of The sequences. RAxML nucleotide by the inferred from support inferred bootstrap Phylogenies Likelihood . 2 FIGURE Maximum
Data from: Assessing the potential of genotyping-by-sequencing-derived single nucleotide polymorphisms to identify the geographic origins of intercepted gypsy moth (Lymantria dispar) specimens: a proof-of-concept study
Open the record for dataset details and reuse information.
Analysis of RNA-seq, DNA target enrichment, and Sanger nucleotide sequence data resolves deep splits in the phylogeny of cuckoo wasps (Hymenoptera: Chrysididae)
Open the record for dataset details and reuse information.
Data from: Restriction site-associated DNA sequencing generates high-quality single nucleotide polymorphisms for assessing hybridization between bighead and silver carp in the United States and China
Open the record for dataset details and reuse information.
Data from: Development of nuclear microsatellite loci and mitochondrial single nucleotide polymorphisms for the natterjack toad, Bufo (Epidalea) calamita (Bufonidae), using next generation sequencing and Competitive Allele Specific PCR (KASPar)
Open the record for dataset details and reuse information.
Data from: Species delimitation of common reef corals in the genus Pocillopora using nucleotide sequence phylogenies, population genetics, and symbiosis ecology
Open the record for dataset details and reuse information.
Data from: A survey of genome-wide single nucleotide polymorphisms through genome re-sequencing in the Périgord black truffle (Tuber melanosporum Vittad.)
Open the record for dataset details and reuse information.
Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages
Open the record for dataset details and reuse information.
Nucleotide sequences in Procambarus clarkii
Open the record for dataset details and reuse information.
Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing
Open the record for dataset details and reuse information.
Data from: Characterization of the transcriptome, nucleotide sequence polymorphism, and natural selection in the desert adapted mouse Peromyscus eremicus
Open the record for dataset details and reuse information.
Model trees and associated simulated nucleotide sequences for testing phylogenetic inference methods
<p>This repository contains 142 tar.gz archive files, each containing nucleotide sequence data that have been simulated using <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> for testing alignment-free phylogenetic inference methods. These datasets were generated by using the results (trees and model parameters) of 142 phylogenomic analyses of real-case data as model (available <a href="https://zenodo.org/record/4034261">here</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>Each archive contains the following files/directories:</p> <ul> <li><code>GTR.params.trees.tsv </code> a tab-delimited file summarizing the real-case GTR+Γ model parameters and the phylogenetic tree used to simulate the sequence dataset (gathered from <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>)</li> <li><code>tax.tsv </code> a tab-delimited file containing the initial (col 1) and simplified (col 2) taxon names</li> <li><code>model.nwk </code> a <a href="https://evolution.genetics.washington.edu/phylip/newicktree.html">Newick</a>-formatted file containing the initial model tree (gathered from <code>GTR.params.trees.tsv</code>) with simplified leaf names (following <code>tax.tsv</code>)</li> <li><code>control.txt </code> the <em>INDELible</em> input file used to simulate the evolution of a sequence along the tree in <code>model.nwk</code></li> <li><code>seq/ </code> a directory containing the simulated sequences (one FASTA file per leaf in the tree in <code>model.nwk</code>)</li> </ul> <p>___</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods
<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+Γ. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+Γ: six rates and one Γ shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p> </p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p> </p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1] </code> integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5] </code> frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10] </code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] </code> Γ shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without Γ) specified to <em>INDELible</em>,</li> <li><code>[12] </code> length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] </code> length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] </code> no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] </code> no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] </code> observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] </code> aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] </code> aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Raw data associated with the article: "Single-molecule DNA sequencing of widely varying GC-content using nucleotide release, capture and detection in microdroplets.", NAR, Puchtler et.al.
<p>All data taken in the production of the corresponding paper: "Single-molecule DNA sequencing of widely varying GC-content using nucleotide release, capture and detection in microdroplets."</p> <p>The associated manuscript describes a method for DNA sequencing which involves the sequential release of nucleotides from a single, immobilised strand of DNA via pyrophosphorolysis (PPL). Released nucleotides, in the form of dNTPs, are captured in microdroplets which are manipulated using an optical-EWOD platform. A detection chemistry within each droplet releases a specific dye depending on which dNTPs are present, allowing the optical read-out of bases within each droplet. Hence, by capturing bases sequentially within droplets as they are cleaved from the strand of DNA, the sequence can be optically identified.</p>
Data from: Mining for single nucleotide polymorphisms and insertions / deletions in expressed sequence tag libraries of oil palm
The oil palm is a tropical oil bearing tree. Recently EST-derived SNPs and SSRs are a free by-product of the currently expanding EST (Expressed Sequence Tag) data bases. The development of high-throughput methods for the detection of SNPs (Single Nucleotide Polymorphism) and small indels (insertion / deletion) has led to a revolution in their use as molecular markers. Available (5452) Oil palm EST sequences were mined from dbEST of NCBI. CAP3 program was used to assemble EST sequences into contigs. Candidate SNPs and Indel polymorphisms were detected using the perl script auto_snip version 1.0 which has used 576 ESTs for detecting SNPs and Indel sites. We found 1180 SNP sites and 137 indel polymorphisms with frequency 1.36 SNPs / 100 bp. Among the six tissues from which the EST libraries had been generated, mesocarp had high frequency of 2.91 SNPs and indels per 100 bp whereas the zygotic embryos had lowest frequency of 0.15 per 100 bp. We also used the Shannon index to analyze the proportion of ten possible types of SNP/indels. ESTs from tissues of normal apex showed highest values of Shannon index (0.60) whereas abnormal apex had least value (0.02). The present report deals the use of Shannon index for comparing SNP/ indel frequencies mined from ESTlibraries and also confirm that the frequency of SNP occurrence in oil palm to use them as markers for genetic studies.
Data from: Development of an Arabis alpina genomic contig sequence dataset and application to single nucleotide polymorphisms discovery
The alpine plant Arabis alpina is an emerging model in the ecological genomic field which is well-suited to identifying the genes involved in local adaptation in contrasted environmental conditions, a subject which remains poorly understood at molecular level. This paper presents the assembly of a pool of A. alpina genomic fragments using Next Generation Sequencing technologies. These contigs cover 172 Mb of the A. alpina genome (i.e. 50% of the genome) and were shown to contain sequences giving positive hits against 96% of the 458 CEGMA core genes (Core Eukaryotic Genes Mapping Approach), a set of highly conserved eukaryotic genes. Regions presenting high nucleic sequence identity with 77% of the close relative Arabidopsis thaliana's genes were found, with an unbiased distribution across the different functional categories of A. thaliana genes. This new resource was tested using a resequencing assay to identify polymorphic sites. Sixteen samples were successfully analyzed and 127,041 Single Nucleotide Polymorphisms identified. This contig dataset will contribute to improving understanding of the ecology of Arabis alpina, thus constituting an important resource for future ecological genomic studies.
TABLE 3. Distance matrix from nucleotide sequences for COI, 16 S in A new species of stauromedusa, Calvadosia festivala (Cnidaria: Staurozoa: Kishinouyeidae) from India
<p><b>TABLE 3.</b> Distance matrix from nucleotide sequences for COI, 16S, and 18S alignments from <i>Calvadosia</i> species closely related to <i>Calvadosia festivala</i> n. sp. “−”: represents absence of data for comparison (COI not available for <i>C. tasmaniensis</i> and <i>C. corbini</i>).</p><table><tbody><tr><th>Species</th><th>1</th><th>2</th><th>3</th><th>4</th><th>5</th></tr></tbody><tbody><tr><th>1. <i>C. tasmaniensis</i></th><td>COI: 0.0000 16S: 0.0000 18S: 0.0000</td><td></td><td></td><td></td><td></td></tr><tr><th>2. <i>C. lewisi</i></th><td>COI: − 16S: 0.0945 18S: 0.0017</td><td>COI: 0.0000 16S: 0.0000 18S: 0.0000</td><td></td><td></td><td></td></tr><tr><th>3. <i>C. corbini</i></th><td>COI: − 16S: 0.0985 18S: 0.0017</td><td>COI: − 16S: 0.1266 18S: 0.0038</td><td>COI: 0.0000 16S: 0.0000 18S: 0.0000</td><td></td><td></td></tr><tr><th>4. <i>C. festivala</i> n. sp.</th><td>COI: − 16S: 0.1400 18S: 0.0000</td><td>COI: 0.2215 16S: 0.1629 18S: 0.0014</td><td>COI: − 16S: 0.1430 18S: 0.0014</td><td>COI: 0.0000 16S: 0.0000 18S: 0.0000</td><td></td></tr><tr><th>5. <i>Calvadosia</i> sp. Moorea</th><td>COI: − 16S: 0.1917 18S: 0.0085</td><td>COI: 0.2506 16S: 0.2197 18S: 0.0070</td><td>COI: − 16S: 0.21009 18S: 0.0045</td><td>COI: 0.2158 16S: 0.2193 18S: 0.0083</td><td>COI: 0.0000 16S: 0.0000 18S: 0.0000</td></tr></tbody></table>
Data from: Single nucleotide polymorphism discovery via genotyping by sequencing to assess population genetic structure and recurrent polyploidization in Andropogon gerardii
Open the record for dataset details and reuse information.
Data from: Mining for single nucleotide polymorphisms and insertions / deletions in expressed sequence tag libraries of oil palm
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.