Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,696

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,696 results for “DNA sequences”

Learn how ShareScore rates datasets ↗
zenodo36/100

Raw DNA sequence data of an indivdual known as "whitequark" (part 2)

<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with 350bp insert. 30-40× coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>

opencc-by-nc-4.0Sep 2017View details →
zenodo36/100

New methods for the genotyping of Legionella pneumophila - Establishment, validation and implementation of a DNA-based microarray and a core genome multilocus sequence typing

<p>This data presented here are part a doctoral thesis with the focus on new genotyping methods for the human pathogen <em>Legionella pneumophila</em>. The data are partially published in articles.&nbsp;</p> <p>The thesis can be downloaded: update of the URL is coming soon</p>

opencc-by-4.0Sep 2013View details →
zenodo36/100

Experimental data for "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads"

<p>The experimental dataset used in "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads."</p> <p>A set of 91,766 150-nt oligos were synthesised with GenScript (oligos.fasta). Each oligo consists of a pseudo-random 110-nt payload flanked by 20-nt primers at each end. The strands are split in three roughly equal groups (two groups of 30,589 and one group of 30,588). Each group has a dedicated primer pair for targeted PCR amplification (the primer pairs used for amplification are provided in primers_synthesis.fasta). The pseudo-random payload was designed to avoid primer-payload collisions.</p> <p>For each file, a sample from the synthesised pool was PCR amplified using the corresponding primer pair and sequenced using Oxford Nanopore Technologies MinION sequencing device following the standard library preparation protocol for amplicon DNA. The raw reads were basecalled using guppy, either in fast- ("acc-false") or high-accuracy ("acc-true") regime. The basecaller generated two groups of reads&mdash;"passQ-true" for the reads that passed the quality-score threshold of 8 and "passQ-false" for those that did not. For each group of reads, a BLAST-based fuzzy search for primer sequences was performed and, based on the resulting alignments, the segments containing the correct primer pairs and located at a distance of 150+-15nt were extracted (separately for forward and reverse-complemented reads). The segments are then assigned to the closest synthesized strand based on Levenshtein distance. The resulting clusters are used to estimate the parameters of the end-to-end DNA storage channel model and to test the proposed error-correction scheme.</p> <p>The archive clustered_read_segments.tar.gz contains 12 sub-archives, for each file (0,1,2), accuracy ("acc-true" or "acc-false"), and Q-score ("passQ-true" or "passQ-false"). Within each sub-archive, there are two folders (one for forward read segments and one for backward read segments), and each folder contains two files: one for the reference synthesised (or "transmitted") sequences that correspond to the file in question ("TX__" &mdash; e.g., "TX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt") and another file for the sequenced (or "received") segment clusters ("RX__" &mdash; e.g., "RX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt"). The received clusters in the "RX__" file are ordered in correspondence with the synthesised sequences in the "TX__" file, and a line "===============================" is used as a separator.</p>

opencc-by-4.0May 2024View details →
dryad36/100

DNA sequences for: Synthetic control of actin polymerization and symmetry breaking in active protocells

<p>Non-linear biomolecular interactions on membranes drive membrane remodeling crucial for biological processes including chemotaxis, cytokinesis, and endocytosis. The complexity of biomolecular interactions, their redundancy, and the importance of spatiotemporal context in membrane organization impede understanding of the physical principles governing membrane mechanics. Developing a minimal in vitro system that mimics molecular signaling and mem- brane remodeling while maintaining physiological fidelity poses a significant challenge. Inspired by chemotaxis, we reconstructed chemically regulated actin polymerization inside vesicles, guiding membrane self-organization. An external, undirected chemical input induced directed actin polymerization and membrane deformation uncorrelated with upstream biochemical cues, suggesting symmetry breaking. A biophysical model incorporating actin dynamics and membrane mechanics proposes that uneven actin distributions cause non-linear membrane deformations, consistent with experimental findings. This protocellular system illuminates the interplay between actin dynamics and membrane shape during symmetry breaking, offering insights into chemotaxis and other cell biological processes.</p>

opencc-zeroMay 2024View details →
dryad36/100

Data from: Affordable de novo generation of fish mitogenomes using amplification-free enrichment of mitochondrial DNA and deep sequencing of long fragments

<p>Biomonitoring surveys from environmental DNA make use of metabarcoding tools to describe the community composition. These studies match their sequencing results against public genomic databases to identify the species. However, mitochondrial genomic reference data are yet incomplete, only a few genes may be available, or the suitability of existing sequence data is suboptimal for species-level resolution. Here we present a dedicated and cost-effective workflow with no DNA amplification for generating complete fish mitogenomes for the purpose of strengthening fish mitochondrial databases. Two different long-fragment sequencing approaches using Oxford Nanopore sequencing coupled with mitochondrial DNA enrichment were used. One where the enrichment is achieved by preferential isolation of mitochondria followed by DNA extraction and nuclear DNA depletion ('mitoenrichment').  A second enrichment approach takes advantage of the CRISPR-Cas9 targeted scission on previously dephosphorylated DNA ('targeted mitosequencing'). The sequencing results varied between tissue, species, and integrity of the DNA. The mitoenrichment method yielded 0.17-12.33 % of sequences on target and a mean coverage ranging from 74.9 to 805-fold. The targeted mitosequencing experiment from native genomic DNA yielded 1.83-55 % of sequences on target and a 38 to 2123-fold mean coverage. This produced complete the mitogenome of species with homopolymeric regions, tandem repeats, and gene rearrangements. We demonstrate that deep sequencing of long fragments of native fish DNA is possible and can be achieved with low computational resources in a cost-effective manner, opening the discovery of mitogenomes of non-model or understudied fish taxa to a broad range of laboratories worldwide.</p>

opencc-zeroJun 2024View details →
zenodo36/100

Fig. 1 in Mitochondrial DNA diversity in the acanthocephalan Prosthenorchis elegans in Colombia based on cytochrome c oxidase I (COI) gene sequence

Fig. 1. Photo showing the characteristic external morphology of Prosthenorchis elegans.

opencc-by-4.0Dec 2015View details →
zenodo36/100

Aligned DNA sequence matrix for phylogenetic analyses of Scinax in the article "Advertisement calls and DNA sequences reveal a new species of Scinax (Anura: Hylidae) on the Pacific Lowlands of Ecuador"

<p>Aligned DNA sequence matrix for phylogenetic analyses of Scinax in the article &quot;Advertisement calls and DNA sequences reveal a new species of Scinax (Anura: Hylidae) on the Pacific Lowlands of Ecuador&quot;. The matrix is in NEXUS format.</p> <p>Gene partition as arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>12S RNA:&nbsp;1-1033</p> <p>16S RNA:&nbsp;1034-2832</p> <p>NADH dehydrogenase subunit 1: 2833-3793</p> <p>Cytochrome Oxidase sub-unit I:&nbsp;3965-4654</p> <p>Cytochrome B:&nbsp;4655-5447</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2017View details →
zenodo36/100

Aligned DNA sequence matrix for phylogenetic analyses of the article "A bizarre new species of Lynchius (Amphibia, Anura, Strabomantidae) from the Andes of Ecuador and first report of Lynchius parkeri in Ecuador"

<p>Aligned DNA sequence matrix for phylogenetic analyses of the article &quot;A bizarre new species of <em>Lynchius</em> (Amphibia, Anura, Strabomantidae) from the Andes of Ecuador and first report of <em>Lynchius parkeri</em> in Ecuador&quot;</p> <p>The matrix is in NEXUS format. Genes are arranged as follows:</p> <p>RAG1:&nbsp;1-652<br> Tyrosinase: 653-1195<br> 12S RNA:&nbsp;1196-2242<br> tRNA Val:&nbsp;2243-2313<br> 16S RNA:&nbsp;2314-3994<br> tRNA Leu = 3995-4065<br> ND1: 4066-5026<br> tRNA Ile: 5027-5144;</p>

opencc-by-4.0Feb 2019View details →
zenodo36/100

Aligned DNA sequence matrix for phylogenetic analyses in the article "Systematics of Huicundomantis, a new subgenus of Pristimantis (Anura, Strabomantidae) with extraordinary cryptic diversity and eleven new species"

<p>Aligned DNA sequence matrix for phylogenetic analyses of the article &quot;Systematics of Huicundomantis, a new subgenus of Pristimantis (Anura, Strabomantidae) with extraordinary cryptic diversity and eleven new species&quot;</p> <p>The matrix is in NEXUS format. Genes are arranged as follows:</p> <p>16S RNA, tRNA-Leu = &nbsp;1&ndash;1408<br> ND1&nbsp;codon position 1 = &nbsp;1409&ndash;2369\3;<br> ND1&nbsp;codon position 2 = &nbsp;1410&ndash;2367\3;<br> ND1&nbsp;codon position 3 = &nbsp;1411&ndash;2368\3;<br> tRNA-Ile, tRNA-Gln, rRNA-Met&nbsp;= 2370&ndash;2544;&nbsp;&nbsp; &nbsp;<br> RAG1&nbsp;codon position 3 = &nbsp;2545&ndash;3199\3;<br> RAG1&nbsp;codon position 1&nbsp;= &nbsp;2546&ndash;3197\3<br> RAG1&nbsp;codon position 2&nbsp;= 2547&ndash;3198\3;</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

Fig. 4. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with ITS2 DNA sequences from 80 in A new genus and three new species of mangrove slugs from the Indo-West Pacific (Mollusca: Gastropoda: Euthyneura: Onchidiidae)

Fig. 4. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with ITS2 DNA sequences from 80 individuals (including 7 outgroups). Numbers by the branches are the bootstrap values (only numbers&gt; 50% are indicated). Numbers for each individual correspond to unique identifiers for DNA extraction. All sequences for specimens of Paromoionchis gen. nov. are new. Information on specimens can be found in the lists of material examined and in Table 1. The letter A corresponds to a clade referred to in the text. The color used for each (mitochondrial) unit is the same as that used in Figs 1–3 and 5–6.

opencc-by-4.0Feb 2019View details →
zenodo36/100

Fig. 3. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with concatenated ITS2 and 28S DNA sequences from 41 in A new genus and three new species of mangrove slugs from the Indo-West Pacific (Mollusca: Gastropoda: Euthyneura: Onchidiidae)

Fig. 3. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with concatenated ITS2 and 28S DNA sequences from 41 individuals (including 7 outgroups). Numbers by the branches are the bootstrap values (only numbers&gt; 50% are indicated). Numbers for each individual correspond to unique identifiers for DNA extraction. All sequences for specimens of Paromoionchis gen. nov. are new. Information on specimens can be found in the lists of material examined and in Table 1. Letters A and B correspond to clades referred to in the text. The color used for each (mitochondrial) unit is the same as that used in Figs 1–2 and 4–6.

opencc-by-4.0Feb 2019View details →
zenodo36/100

Aligned DNA sequence matrixes for the study of the divergent times of phytoplasmas

<p>Sequence alignments of 16S rRNA and&nbsp;methionine aminopeptidase (map) are provided in FASTA files &ldquo;Cao_et_al_16S.fas&rdquo; and &ldquo;Cao_et_al_map.fas&rdquo;, respectively. Detailed information of the data matrixes is as follows:</p> <p>&nbsp;</p> <p>File name: Cao_et_al_16S.fas</p> <p>Number of taxa: 220</p> <p>Number of characters: 1655</p> <p>Gap: -</p> <p>&nbsp;</p> <p>File name: Cao_et_al_map.fas</p> <p>Number of taxa: 83</p> <p>Number of characters: 564</p> <p>Gap: -</p>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Aligned DNA sequence matrixes for the study of the divergent times of phytoplasmas

<p>Sequence alignments of 16S rRNA and&nbsp;methionine aminopeptidase (map) are provided in FASTA files &ldquo;Cao_et_al_16S.fas&rdquo; and &ldquo;Cao_et_al_map.fas&rdquo;, respectively. Detailed information of the data matrixes is as follows:</p> <p>&nbsp;</p> <p>File name: Cao_et_al_16S.fas</p> <p>Number of taxa: 220</p> <p>Number of characters: 1655</p> <p>Gap: -</p> <p>&nbsp;</p> <p>File name: Cao_et_al_map.fas</p> <p>Number of taxa: 83</p> <p>Number of characters: 564</p> <p>Gap: -</p>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Genus level DNA sequence data for three genes (matK, rbcL, trnH-psbA) for the paper: A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability

<p>This data file contains the consensus DNA sequences in fasta format, of 64 tree genera found in Mediterranean Europe, following the checklist of M&eacute;dail et al. (2019).&nbsp;</p> <p>The data are used in a manuscript submitted for publication to Botany Letters and currently under revision. The manuscript is entitled: &quot;<em>A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability</em>&quot;. Its authors are: Marwan Cheikh Albassatneh, Marcial Escudero, Loic Ponge<sup>*</sup>, Anne-Christine Monnet, Juan Arroyo, Toni Nikolic, Gianluigi Bacchetta, Francesca Bagnoli, Panayotis Dimopoulos, Agathe Leriche, Fr&eacute;d&eacute;ric M&eacute;dail, Anne Roig, Ilaria Spanu, Giovanni Giuseppe Vendramin, Arndt Hampe, Bruno Fady.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Aligned DNA sequence matrix for phylogenetic analyses in the article "A new glassfrog of the genus Nymphargus (Anura: Centrolenidae) from Cordillera del Cóndor, Ecuador"

<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "<span>A new glassfrog of the genus Nymphargus</span> (Anura<span>:</span> Centrolenidae) <span>from Cordillera del C&oacute;ndor, Ecuador</span>"</p> <p>The matrix is in NEXUS format and has 6613 bp and 102 terminals.</p> <p>Partitions are as follows:</p> <div> <div>charset 12S = 1-944;</div> <div>charset 16S = 945-2093;</div> <div>charset BNDFcodonPos1 = &nbsp;2096-2792\3;</div> <div>charset BNDFcodonPos2 = &nbsp;2094-2793\3;</div> <div>charset BNDFcodonPos3 = &nbsp;2095-2791\3;</div> <div>charset ND1codonPos1 = &nbsp;2795-3749\3;</div> <div>charset ND1codonPos2 = &nbsp;2796-3750\3;</div> <div>charset ND1codonPos3 = &nbsp;2794-3751\3;</div> <div>charset CXCR4codonPos1 = &nbsp;3753-4107\3;</div> <div>charset CXCR4codonPos2 = &nbsp;3754-4105\3;</div> <div>charset CXCR4codonPos3 = &nbsp;3752-4106\3;</div> <div>charset cmyccodonPos1 = &nbsp;4108-4507\3;</div> <div>charset cmyccodonPos2 = &nbsp;4109-4508\3;</div> <div>charset cmyccodonPos3 = &nbsp;4110-4509\3;</div> <div>charset POMCcodonPos1 = &nbsp;4511-5120\3;</div> <div>charset POMCcodonPos2 = &nbsp;4512-5121\3;</div> <div>charset POMCcodonPos3 = &nbsp;4510-5122\3;</div> <div>charset RAG1codonPos1 = &nbsp;5123-5576\3;</div> <div>charset RAG1codonPos2 = &nbsp;5124-5577\3;</div> <div>charset RAG1codonPos3 = &nbsp;5125-5578\3;</div> <div>charset SLC8A1codonPos1 = &nbsp;5580-6120\3;</div> <div>charset SLC8A1codonPos2 = &nbsp;5581-6118\3;</div> <div>charset SLC8A1codonPos3 = &nbsp;5579-6119\3;</div> <div>charset SLC8A3codonPos1 = &nbsp;6122-6587\3;</div> <div>charset SLC8A3codonPos2 = &nbsp;6123-6585\3;</div> <div>charset SLC8A3codonPos3 = &nbsp;6121-6586\3;</div> </div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data for paper: Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences

<p>This deposit contains data for the paper entitled: "<strong>Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences</strong>"</p> <p>Contents include:</p> <ul> <li><strong>00_sequencing-metadata.xlsx</strong> - Metadata for each sequencing sample.</li> <li><strong>01_raw-fastqs.zip</strong> - Raw basecalled FASTQ data.</li> <li><strong>02_clean-fastas.zip</strong> - Cleaned FASTA files (removal of adapters and barcode sequences).</li> <li><strong>03_read-statistics.zip</strong> - Read statistics for all samples.</li> <li><strong>04_sequence-analysis.zip</strong> - Sequence analysis output for all samples.</li> <li><strong>05_qpcr-data.xlsx</strong> - qPCR data for all the dNTP mixes studied.</li> <li><strong>06_analysis-scripts.zip</strong> - Analysis scripts.</li> </ul>

opencc-by-4.0Aug 2024View details →
dryad36/100

Data from: How "simple" methodological decisions affect interpretation of population structure based on reduced representation library DNA sequencing: a case study using the lake whitefish

Reduced representation (RRL) sequencing approaches (e.g., RADSeq, genotyping by sequencing) require decisions about how much to invest in genome coverage and sequencing depth (library quality), as well as choices of values for adjustable bioinformatics parameters. To empirically explore the importance of these "simple" decisions, we generated two independent sequencing libraries for the same 142 individual lake whitefish (Coregonus clupeaformis) using a nextRAD RRL approach: (1) A small number of loci and low sequencing depth (library A); and (2) more loci and higher sequencing depth (library B). The fish were selected from populations with different levels of expected genetic subdivision. Each library was analyzed using the STACKS pipeline followed by three types of population structure assessment (FST, DAPC and ADMIXTURE) with iterative increases in the stringency of sequencing depth and missing data requirements, as well as more specific a priori population maps. Library B was always able to resolve strong population differentiation in all three types of assessment regardless of the selected parameters. In contrast, library A produced more variable results; increasing the minimum sequencing depth threshold (-m) resulted in a reduced number of retained loci, and therefore lost resolution at high -m values for FST and ADMIXTURE, but not DAPC. FST and DAPC were robust to varying the population map and increasing the stringency of missing data requirements. In contrast, ADMIXTURE was unable to resolve strong population differentiation when increasing these same parameters in library A. Similarly, when examining fine scale population subdivision, library B was robust to changing parameters but library A lost resolution depending on the parameter set. We used library B to examine actual subdivision in our study populations. All three types of analysis found complete subdivision among populations in Lake Huron, ON and Dore Lake, SK, Canada using 10,640 SNP loci. Weak population subdivision was detected in Lake Huron with fish from sites in the north-west, Search Bay, North Point and Hammond Bay, showing slight differentiation. Overall, we show that apparently simple decisions about library quality and bioinformatics parameters can have potentially important impacts on the interpretation of population subdivision. Although costly, the early investment in a high-quality library and more conservative stringency settings on STACKS parameters lead to a final dataset that was more consistent and robust when examining both weak and strong population differentiation.

opencc-zeroMar 2020View details →
dryad36/100

DNA sequences of transgenes detected via environmental DNA (raw ABI files, processed FASTA files, and reference alignments)

We demonstrate that simple, non-invasive environmental DNA (eDNA) methods can detect transgenes of genetically modified (GM) animals from terrestrial and aquatic sources in invertebrate and vertebrate systems. We detected transgenic fragments between 82-234 bp through targeted PCR amplification of environmental DNA extracted from food media of GM fruit flies (<i>Drosophila melanogaster</i>), feces, urine, and saliva of GM laboratory mice (<i>Mus musculus</i>), and aquarium water of GM tetra fish (<i>Gymnocorymbus ternetzi</i>). With rapidly growing accessibility of genome-editing technologies such as CRISPR, the prevalence and diversity of GM animals will increase dramatically. GM animals have already been released into the wild with more releases planned in the future. eDNA methods have the potential to address the critical need for sensitive, accurate, and cost-effective detection and monitoring of GM animals and their transgenes in nature.

opencc-zeroAug 2021View details →
zenodo36/100

Fig. 5 in Copelatus sibelaemontis sp. nov. (Coleoptera: Dytiscidae) from the Moluccas with generic assignment based on morphology and DNA sequence data

Fig. 5. Distribution of Copelatus sibelaemontis sp. nov.

opencc-by-4.0Dec 2010View details →
zenodo36/100

FIGURE 2 in Voucher Specimens For Dna Sequences Of Phytoseiid Mites (Acari: Mesostigmata)

FIGURE 2: DNA bands of the expected sizes were amplified for the mitochondrial Cytb fragment from A – Neoseiulus idaeus, from B – from Typhlodromus (Typhlodromus) exhilaratus. A – Lane M shows the molecular weight marker, C indicates the water control, C+ indicates the positive control, lanes 1-6 show the PCR products of a single N. idaeus specimen preliminarily treated with lactic acid. Image B: Lane M shows the molecular weight marker, Lanes 1-5 show the PCR products of a single T. (T.) exhilaratus specimen conserved in 100 % alcohol.

opencc-by-nd-4.0Dec 2010View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record