Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,696
datasets available to search
ShareScore release 0.9.0
Dataset results
1,696 results for “DNA sequences”
Raw DNA sequence data of an indivdual known as "whitequark" (part 2)
<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with 350bp insert. 30-40× coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>
New methods for the genotyping of Legionella pneumophila - Establishment, validation and implementation of a DNA-based microarray and a core genome multilocus sequence typing
<p>This data presented here are part a doctoral thesis with the focus on new genotyping methods for the human pathogen <em>Legionella pneumophila</em>. The data are partially published in articles. </p> <p>The thesis can be downloaded: update of the URL is coming soon</p>
Experimental data for "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads"
<p>The experimental dataset used in "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads."</p> <p>A set of 91,766 150-nt oligos were synthesised with GenScript (oligos.fasta). Each oligo consists of a pseudo-random 110-nt payload flanked by 20-nt primers at each end. The strands are split in three roughly equal groups (two groups of 30,589 and one group of 30,588). Each group has a dedicated primer pair for targeted PCR amplification (the primer pairs used for amplification are provided in primers_synthesis.fasta). The pseudo-random payload was designed to avoid primer-payload collisions.</p> <p>For each file, a sample from the synthesised pool was PCR amplified using the corresponding primer pair and sequenced using Oxford Nanopore Technologies MinION sequencing device following the standard library preparation protocol for amplicon DNA. The raw reads were basecalled using guppy, either in fast- ("acc-false") or high-accuracy ("acc-true") regime. The basecaller generated two groups of reads—"passQ-true" for the reads that passed the quality-score threshold of 8 and "passQ-false" for those that did not. For each group of reads, a BLAST-based fuzzy search for primer sequences was performed and, based on the resulting alignments, the segments containing the correct primer pairs and located at a distance of 150+-15nt were extracted (separately for forward and reverse-complemented reads). The segments are then assigned to the closest synthesized strand based on Levenshtein distance. The resulting clusters are used to estimate the parameters of the end-to-end DNA storage channel model and to test the proposed error-correction scheme.</p> <p>The archive clustered_read_segments.tar.gz contains 12 sub-archives, for each file (0,1,2), accuracy ("acc-true" or "acc-false"), and Q-score ("passQ-true" or "passQ-false"). Within each sub-archive, there are two folders (one for forward read segments and one for backward read segments), and each folder contains two files: one for the reference synthesised (or "transmitted") sequences that correspond to the file in question ("TX__" — e.g., "TX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt") and another file for the sequenced (or "received") segment clusters ("RX__" — e.g., "RX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt"). The received clusters in the "RX__" file are ordered in correspondence with the synthesised sequences in the "TX__" file, and a line "===============================" is used as a separator.</p>
DNA sequences for: Synthetic control of actin polymerization and symmetry breaking in active protocells
<p>Non-linear biomolecular interactions on membranes drive membrane remodeling crucial for biological processes including chemotaxis, cytokinesis, and endocytosis. The complexity of biomolecular interactions, their redundancy, and the importance of spatiotemporal context in membrane organization impede understanding of the physical principles governing membrane mechanics. Developing a minimal in vitro system that mimics molecular signaling and mem- brane remodeling while maintaining physiological fidelity poses a significant challenge. Inspired by chemotaxis, we reconstructed chemically regulated actin polymerization inside vesicles, guiding membrane self-organization. An external, undirected chemical input induced directed actin polymerization and membrane deformation uncorrelated with upstream biochemical cues, suggesting symmetry breaking. A biophysical model incorporating actin dynamics and membrane mechanics proposes that uneven actin distributions cause non-linear membrane deformations, consistent with experimental findings. This protocellular system illuminates the interplay between actin dynamics and membrane shape during symmetry breaking, offering insights into chemotaxis and other cell biological processes.</p>
Data from: Affordable de novo generation of fish mitogenomes using amplification-free enrichment of mitochondrial DNA and deep sequencing of long fragments
<p>Biomonitoring surveys from environmental DNA make use of metabarcoding tools to describe the community composition. These studies match their sequencing results against public genomic databases to identify the species. However, mitochondrial genomic reference data are yet incomplete, only a few genes may be available, or the suitability of existing sequence data is suboptimal for species-level resolution. Here we present a dedicated and cost-effective workflow with no DNA amplification for generating complete fish mitogenomes for the purpose of strengthening fish mitochondrial databases. Two different long-fragment sequencing approaches using Oxford Nanopore sequencing coupled with mitochondrial DNA enrichment were used. One where the enrichment is achieved by preferential isolation of mitochondria followed by DNA extraction and nuclear DNA depletion ('mitoenrichment'). A second enrichment approach takes advantage of the CRISPR-Cas9 targeted scission on previously dephosphorylated DNA ('targeted mitosequencing'). The sequencing results varied between tissue, species, and integrity of the DNA. The mitoenrichment method yielded 0.17-12.33 % of sequences on target and a mean coverage ranging from 74.9 to 805-fold. The targeted mitosequencing experiment from native genomic DNA yielded 1.83-55 % of sequences on target and a 38 to 2123-fold mean coverage. This produced complete the mitogenome of species with homopolymeric regions, tandem repeats, and gene rearrangements. We demonstrate that deep sequencing of long fragments of native fish DNA is possible and can be achieved with low computational resources in a cost-effective manner, opening the discovery of mitogenomes of non-model or understudied fish taxa to a broad range of laboratories worldwide.</p>
Fig. 1 in Mitochondrial DNA diversity in the acanthocephalan Prosthenorchis elegans in Colombia based on cytochrome c oxidase I (COI) gene sequence
Fig. 1. Photo showing the characteristic external morphology of Prosthenorchis elegans.
Aligned DNA sequence matrix for phylogenetic analyses of Scinax in the article "Advertisement calls and DNA sequences reveal a new species of Scinax (Anura: Hylidae) on the Pacific Lowlands of Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of Scinax in the article "Advertisement calls and DNA sequences reveal a new species of Scinax (Anura: Hylidae) on the Pacific Lowlands of Ecuador". The matrix is in NEXUS format.</p> <p>Gene partition as arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>12S RNA: 1-1033</p> <p>16S RNA: 1034-2832</p> <p>NADH dehydrogenase subunit 1: 2833-3793</p> <p>Cytochrome Oxidase sub-unit I: 3965-4654</p> <p>Cytochrome B: 4655-5447</p> <p> </p>
Aligned DNA sequence matrix for phylogenetic analyses of the article "A bizarre new species of Lynchius (Amphibia, Anura, Strabomantidae) from the Andes of Ecuador and first report of Lynchius parkeri in Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "A bizarre new species of <em>Lynchius</em> (Amphibia, Anura, Strabomantidae) from the Andes of Ecuador and first report of <em>Lynchius parkeri</em> in Ecuador"</p> <p>The matrix is in NEXUS format. Genes are arranged as follows:</p> <p>RAG1: 1-652<br> Tyrosinase: 653-1195<br> 12S RNA: 1196-2242<br> tRNA Val: 2243-2313<br> 16S RNA: 2314-3994<br> tRNA Leu = 3995-4065<br> ND1: 4066-5026<br> tRNA Ile: 5027-5144;</p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Systematics of Huicundomantis, a new subgenus of Pristimantis (Anura, Strabomantidae) with extraordinary cryptic diversity and eleven new species"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "Systematics of Huicundomantis, a new subgenus of Pristimantis (Anura, Strabomantidae) with extraordinary cryptic diversity and eleven new species"</p> <p>The matrix is in NEXUS format. Genes are arranged as follows:</p> <p>16S RNA, tRNA-Leu = 1–1408<br> ND1 codon position 1 = 1409–2369\3;<br> ND1 codon position 2 = 1410–2367\3;<br> ND1 codon position 3 = 1411–2368\3;<br> tRNA-Ile, tRNA-Gln, rRNA-Met = 2370–2544; <br> RAG1 codon position 3 = 2545–3199\3;<br> RAG1 codon position 1 = 2546–3197\3<br> RAG1 codon position 2 = 2547–3198\3;</p>
Fig. 4. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with ITS2 DNA sequences from 80 in A new genus and three new species of mangrove slugs from the Indo-West Pacific (Mollusca: Gastropoda: Euthyneura: Onchidiidae)
Fig. 4. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with ITS2 DNA sequences from 80 individuals (including 7 outgroups). Numbers by the branches are the bootstrap values (only numbers> 50% are indicated). Numbers for each individual correspond to unique identifiers for DNA extraction. All sequences for specimens of Paromoionchis gen. nov. are new. Information on specimens can be found in the lists of material examined and in Table 1. The letter A corresponds to a clade referred to in the text. The color used for each (mitochondrial) unit is the same as that used in Figs 1–3 and 5–6.
Fig. 3. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with concatenated ITS2 and 28S DNA sequences from 41 in A new genus and three new species of mangrove slugs from the Indo-West Pacific (Mollusca: Gastropoda: Euthyneura: Onchidiidae)
Fig. 3. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with concatenated ITS2 and 28S DNA sequences from 41 individuals (including 7 outgroups). Numbers by the branches are the bootstrap values (only numbers> 50% are indicated). Numbers for each individual correspond to unique identifiers for DNA extraction. All sequences for specimens of Paromoionchis gen. nov. are new. Information on specimens can be found in the lists of material examined and in Table 1. Letters A and B correspond to clades referred to in the text. The color used for each (mitochondrial) unit is the same as that used in Figs 1–2 and 4–6.
Aligned DNA sequence matrixes for the study of the divergent times of phytoplasmas
<p>Sequence alignments of 16S rRNA and methionine aminopeptidase (map) are provided in FASTA files “Cao_et_al_16S.fas” and “Cao_et_al_map.fas”, respectively. Detailed information of the data matrixes is as follows:</p> <p> </p> <p>File name: Cao_et_al_16S.fas</p> <p>Number of taxa: 220</p> <p>Number of characters: 1655</p> <p>Gap: -</p> <p> </p> <p>File name: Cao_et_al_map.fas</p> <p>Number of taxa: 83</p> <p>Number of characters: 564</p> <p>Gap: -</p>
Aligned DNA sequence matrixes for the study of the divergent times of phytoplasmas
<p>Sequence alignments of 16S rRNA and methionine aminopeptidase (map) are provided in FASTA files “Cao_et_al_16S.fas” and “Cao_et_al_map.fas”, respectively. Detailed information of the data matrixes is as follows:</p> <p> </p> <p>File name: Cao_et_al_16S.fas</p> <p>Number of taxa: 220</p> <p>Number of characters: 1655</p> <p>Gap: -</p> <p> </p> <p>File name: Cao_et_al_map.fas</p> <p>Number of taxa: 83</p> <p>Number of characters: 564</p> <p>Gap: -</p>
Genus level DNA sequence data for three genes (matK, rbcL, trnH-psbA) for the paper: A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability
<p>This data file contains the consensus DNA sequences in fasta format, of 64 tree genera found in Mediterranean Europe, following the checklist of Médail et al. (2019). </p> <p>The data are used in a manuscript submitted for publication to Botany Letters and currently under revision. The manuscript is entitled: "<em>A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability</em>". Its authors are: Marwan Cheikh Albassatneh, Marcial Escudero, Loic Ponge<sup>*</sup>, Anne-Christine Monnet, Juan Arroyo, Toni Nikolic, Gianluigi Bacchetta, Francesca Bagnoli, Panayotis Dimopoulos, Agathe Leriche, Frédéric Médail, Anne Roig, Ilaria Spanu, Giovanni Giuseppe Vendramin, Arndt Hampe, Bruno Fady.</p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "A new glassfrog of the genus Nymphargus (Anura: Centrolenidae) from Cordillera del Cóndor, Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "<span>A new glassfrog of the genus Nymphargus</span> (Anura<span>:</span> Centrolenidae) <span>from Cordillera del Cóndor, Ecuador</span>"</p> <p>The matrix is in NEXUS format and has 6613 bp and 102 terminals.</p> <p>Partitions are as follows:</p> <div> <div>charset 12S = 1-944;</div> <div>charset 16S = 945-2093;</div> <div>charset BNDFcodonPos1 = 2096-2792\3;</div> <div>charset BNDFcodonPos2 = 2094-2793\3;</div> <div>charset BNDFcodonPos3 = 2095-2791\3;</div> <div>charset ND1codonPos1 = 2795-3749\3;</div> <div>charset ND1codonPos2 = 2796-3750\3;</div> <div>charset ND1codonPos3 = 2794-3751\3;</div> <div>charset CXCR4codonPos1 = 3753-4107\3;</div> <div>charset CXCR4codonPos2 = 3754-4105\3;</div> <div>charset CXCR4codonPos3 = 3752-4106\3;</div> <div>charset cmyccodonPos1 = 4108-4507\3;</div> <div>charset cmyccodonPos2 = 4109-4508\3;</div> <div>charset cmyccodonPos3 = 4110-4509\3;</div> <div>charset POMCcodonPos1 = 4511-5120\3;</div> <div>charset POMCcodonPos2 = 4512-5121\3;</div> <div>charset POMCcodonPos3 = 4510-5122\3;</div> <div>charset RAG1codonPos1 = 5123-5576\3;</div> <div>charset RAG1codonPos2 = 5124-5577\3;</div> <div>charset RAG1codonPos3 = 5125-5578\3;</div> <div>charset SLC8A1codonPos1 = 5580-6120\3;</div> <div>charset SLC8A1codonPos2 = 5581-6118\3;</div> <div>charset SLC8A1codonPos3 = 5579-6119\3;</div> <div>charset SLC8A3codonPos1 = 6122-6587\3;</div> <div>charset SLC8A3codonPos2 = 6123-6585\3;</div> <div>charset SLC8A3codonPos3 = 6121-6586\3;</div> </div>
Data for paper: Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences
<p>This deposit contains data for the paper entitled: "<strong>Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences</strong>"</p> <p>Contents include:</p> <ul> <li><strong>00_sequencing-metadata.xlsx</strong> - Metadata for each sequencing sample.</li> <li><strong>01_raw-fastqs.zip</strong> - Raw basecalled FASTQ data.</li> <li><strong>02_clean-fastas.zip</strong> - Cleaned FASTA files (removal of adapters and barcode sequences).</li> <li><strong>03_read-statistics.zip</strong> - Read statistics for all samples.</li> <li><strong>04_sequence-analysis.zip</strong> - Sequence analysis output for all samples.</li> <li><strong>05_qpcr-data.xlsx</strong> - qPCR data for all the dNTP mixes studied.</li> <li><strong>06_analysis-scripts.zip</strong> - Analysis scripts.</li> </ul>
Data from: How "simple" methodological decisions affect interpretation of population structure based on reduced representation library DNA sequencing: a case study using the lake whitefish
Reduced representation (RRL) sequencing approaches (e.g., RADSeq, genotyping by sequencing) require decisions about how much to invest in genome coverage and sequencing depth (library quality), as well as choices of values for adjustable bioinformatics parameters. To empirically explore the importance of these "simple" decisions, we generated two independent sequencing libraries for the same 142 individual lake whitefish (Coregonus clupeaformis) using a nextRAD RRL approach: (1) A small number of loci and low sequencing depth (library A); and (2) more loci and higher sequencing depth (library B). The fish were selected from populations with different levels of expected genetic subdivision. Each library was analyzed using the STACKS pipeline followed by three types of population structure assessment (FST, DAPC and ADMIXTURE) with iterative increases in the stringency of sequencing depth and missing data requirements, as well as more specific a priori population maps. Library B was always able to resolve strong population differentiation in all three types of assessment regardless of the selected parameters. In contrast, library A produced more variable results; increasing the minimum sequencing depth threshold (-m) resulted in a reduced number of retained loci, and therefore lost resolution at high -m values for FST and ADMIXTURE, but not DAPC. FST and DAPC were robust to varying the population map and increasing the stringency of missing data requirements. In contrast, ADMIXTURE was unable to resolve strong population differentiation when increasing these same parameters in library A. Similarly, when examining fine scale population subdivision, library B was robust to changing parameters but library A lost resolution depending on the parameter set. We used library B to examine actual subdivision in our study populations. All three types of analysis found complete subdivision among populations in Lake Huron, ON and Dore Lake, SK, Canada using 10,640 SNP loci. Weak population subdivision was detected in Lake Huron with fish from sites in the north-west, Search Bay, North Point and Hammond Bay, showing slight differentiation. Overall, we show that apparently simple decisions about library quality and bioinformatics parameters can have potentially important impacts on the interpretation of population subdivision. Although costly, the early investment in a high-quality library and more conservative stringency settings on STACKS parameters lead to a final dataset that was more consistent and robust when examining both weak and strong population differentiation.
DNA sequences of transgenes detected via environmental DNA (raw ABI files, processed FASTA files, and reference alignments)
We demonstrate that simple, non-invasive environmental DNA (eDNA) methods can detect transgenes of genetically modified (GM) animals from terrestrial and aquatic sources in invertebrate and vertebrate systems. We detected transgenic fragments between 82-234 bp through targeted PCR amplification of environmental DNA extracted from food media of GM fruit flies (<i>Drosophila melanogaster</i>), feces, urine, and saliva of GM laboratory mice (<i>Mus musculus</i>), and aquarium water of GM tetra fish (<i>Gymnocorymbus ternetzi</i>). With rapidly growing accessibility of genome-editing technologies such as CRISPR, the prevalence and diversity of GM animals will increase dramatically. GM animals have already been released into the wild with more releases planned in the future. eDNA methods have the potential to address the critical need for sensitive, accurate, and cost-effective detection and monitoring of GM animals and their transgenes in nature.
Fig. 5 in Copelatus sibelaemontis sp. nov. (Coleoptera: Dytiscidae) from the Moluccas with generic assignment based on morphology and DNA sequence data
Fig. 5. Distribution of Copelatus sibelaemontis sp. nov.
FIGURE 2 in Voucher Specimens For Dna Sequences Of Phytoseiid Mites (Acari: Mesostigmata)
FIGURE 2: DNA bands of the expected sizes were amplified for the mitochondrial Cytb fragment from A – Neoseiulus idaeus, from B – from Typhlodromus (Typhlodromus) exhilaratus. A – Lane M shows the molecular weight marker, C indicates the water control, C+ indicates the positive control, lanes 1-6 show the PCR products of a single N. idaeus specimen preliminarily treated with lactic acid. Image B: Lane M shows the molecular weight marker, Lanes 1-5 show the PCR products of a single T. (T.) exhilaratus specimen conserved in 100 % alcohol.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.