Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
36
datasets available to search
ShareScore release 0.7.1
Dataset results
36 results for “Oxford Nanopore”
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
Chenopodium Oxford Nanopore libraries
<p>Oxford Nanopore libraries of four diploid Chenopodium species (2n=2x=18), <em>C. acuminatum</em>, <em>C. iljinii</em>, <em>C. pamiricum</em> and <em>C. suecicum. </em>The 50,000 longest ON reads in each library. </p>
DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing
<p>Curated dataset for the manuscript named "DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing".</p> <p>Five publicly available datasets sequenced on ONT MinION/GridION were used for the experiments (see below for original sources). These datasets contained raw signal data in single-FAST5 format (one file per each read), which were converted to BLOW5 format using slow5tools to enable convenient and efficient file manipulation. Then, 40,000 reads containing at least 4500 signal samples were extracted from each dataset. From each dataset, 20,000 reads are for training (<species>/train-<species>.blow5) and the rest for testing (<species>/test-<species>.blow5). Basecalled reads for the dataset used for testing are also available (test-<species>.fastq). Guppy version 6.1.3 under dna_r9.4.1_450bps_hac mode was used. The reference genomes are also given (<species>/<species>-ref.fasta)</p> <p>Original datasets are from the following sources:<br> SARS-CoV-2: https://community.artic.network/t/links-to-raw-fast5-fastq-data-for-artic-protocol/17<br> Zymo Metagenome: https://github.com/LomanLab/mockcommunity<br> Chlamydomonas: https://sra-download.ncbi.nlm.nih.gov/traces/era20/ERZ/003237/ERR3237140/Chlamydomonas_0.tar.gz<br> Saccharomyces cerevisiae: https://www.ncbi.nlm.nih.gov/bioproject/PRJNA510813</p>
Hieracium alpinun PAI33838 (2n = 2x = 18) Oxford Nanopore Technology sequences library
<p>Sample 1 000 000 reads (trimmed).</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
<p>In the paper associated with this dataset, a workflow is presented which enables the identification of single-/low-copy nuclear molecular markers for a plant group of interest, by mining data from a small representative target capture experiment done using a commercial probe kit and Oxford Nanopore long-read sequencing. The proposed pipeline first assesses sequence variability contained in the data from targeted loci and assigns reads to their respective genes, via a combined BLAST/clustering procedure. Cluster consensus sequences are then examined based on four pre-defined criteria presumably indicative for absence of paralogy. This is done by calculating four specialized indices; loci are ranked according to their performance in these indices, and top-scoring loci are considered putatively single- or low-copy. The approach can be applied to any probe set. As it relies on long reads, the contribution also provides template workflows for processing Nanopore-based target capture data. Identified loci can be used for NGS amplicon sequencing. For detection of possibly remaining paralogy in these data, which might occur in groups with rampant paralogy, the long-read assembly tool CANU is employed. The presented workflow can be useful for researchers dealing with reticulate or polyploidization phylogenetic histories in plants.</p> <p>The present dataset contains several documents supplementing the original paper. Its most important elements are a detailed description (alongside two graphical workflow figures) of all methods employed in the study, suitable for reproducing the steps of the workflow and also the wet-lab work. The workflow employs a collection of BASH, Python and R scripts which is available here, together with a detailed account on command line use in Linux. Also, reference sequences for the identified markers can be found as well as sequence alignments derived from the amplicon sequencing.</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
Open the record for dataset details and reuse information.
Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing
<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>
Oxford Nanopore sequencing for comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity
<p><span>The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here, we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. </span></span></p>
Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns (repository for Genome Research paper, 2022)
<p>Simulated ONT and PacBio RNA-Seq data for "Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns" paper (Mikheenko et al., Genome Research, 2022). All details can be found in the Methods section of the paper.</p> <p><strong>PacBio.simulated_uniform_coverage.fasta.gz</strong> and <strong>ONT.simulated_uniform_coverage.fasta.gz files</strong> were used in Supplemental Note “Benchmarking of the read-to-isoform assignment algorithm”.</p> <p><strong>ONT.simulated_real_expression.fasta.gz</strong> file and all GTF files were used in the Section "Splice site correction improves transcript discovery precision". <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> was used as the annotation file for all tools. <strong>mouse.gencode.M26.spatial.15percent.expressed.gtf </strong>contains the set of all expressed isoforms. <strong>mouse.gencode.M26.spatial.15percent.expressed_kept.gtf</strong> contains those of the isoforms that are in presented in the annotation file ("known" transcripts), <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> contains expressed isoforms that were removed from the annotation ("novel" transcripts).</p>
Rapid and Inexpensive Whole-Genome Sequencing of SARS-CoV2 using 1200 bp Tiled Amplicons and Oxford Nanopore Rapid Barcoding
<p>Description of 1200bp amplicon primer sets and .bed and .tsv files for SARS-CoV-2 assembly using the ARTIC bioinformatics pipeline.</p>
Oxford Nanopore Direct RNA Sequencing datasets for detecting rRNA modifications in the Brassica oleracea mitoribosome
<p>Oxford Nanopore Direct RNA Sequencing (DRS) was applied for the detection of rRNA modifications in the <em>Brassica oleracea</em> mitoribosome. A comparison between native rRNA transcripts and in vitro transcribed (IVT) rRNA transcripts that were devoid of any modification indicated systematic base-calling errors and/or variations in current intensities that led to the prediction of the modified nucleotides (Begik et al., 2018).</p> <p>For Nanopore (DRS) library preparation, custom reverse transcription adapters (RTAs) containing Deeplexicon multiplexing barcodes (BC1, BC2 or BC3) were designed for sequence-specific ligation to the 3’-ends of <em>B. oleracea</em> mitochondrial rRNAs.</p> <p>Basecalling and demultiplexing of ONT direct RNA sequencing data were performed by Guppy and Deeplexicon, respectively.</p> <p>For the analysis and visualization of current intensities, a full description is available in the associated publication.</p> <p>The <strong>Eventalign_18S.R</strong> Rscript was used to make the nanopore <strong>18S </strong>signal analysis</p> <p>The <strong>Eventalign_26S.R</strong> Rscript was used to make the nanopore <strong>26S </strong>signal analysis</p> <p>The raw data for the scripts are in the following folders:</p> <p>- For 26S: 26S_Native.barcode01 and 26S_IVT.barcode03<br>- For 18S: 18S_Native.barcode01 and 18S_IVT.barcode02</p> <p>These folders contain the read alignment files output from minimap2 (in .bam format) and the eventalign files generated by the f5c software (in tsv format).</p> <p>The FASTA folder contains the mitochondrial 18S and 26S rRNA gene reference sequence in fasta format</p>
A near telomere-to-telomere phased reference assembly for the male mountain gorilla (Gorilla beringei beringei) -Oxford Nanopore reads
<p>The critically endangered mountain gorilla <em>Gorilla beringei beringei</em> faces numerous threats to its survival, highlighting the urgent need for genomic resources to aid conservation efforts. Here, we present a near telomere-to-telomere, haplotype-phased reference genome assembly for a male mountain gorilla generated using Pacbio HiFi and Oxford Nanopore Ultralong data. The resulting assembly exhibits exceptional contiguity, with contig N50 of ~ 95 Mbps for the combined pseudohaplotype (3,540,458,497 bps, and 56.5 Mbps (3.1 Gbps) and 51.0 Mbps (3.2 Gbps) for the maternal and paternal haplotypes and an average QV of 65.15 (error rate = 3.1 x 10-7) and 0% switch errors detected. These represent substantial improvements over most other available primate genomes. This high-quality reference genome provides an invaluable resource for future studies on gorilla evolution, adaptation, and conservation, ultimately contributing to the long-term survival of this iconic species.</p> <p>This repository hosts the Oxford Nanopore ultralong reads.</p> <p>The preprint for this work is available at bioRXiv, doi: https://doi.org/10.1101/2024.10.28.620258. </p>
Matched Oxford Nanopore Technologies and Bisulfite Sequencing of the GM24385 Cell Line
<p>One of the most widespread genomic modifications is 5-methylcytosine (5mC), which most frequently occurs at <a href="https://en.wikipedia.org/wiki/CpG_site">CpG</a> dinucleotides. Compared to whole-genome bisulfite sequencing, the traditional method of 5mC detection, nanopore technology can offer many advantages such as simplicity of sample prep and subsequent analysis.</p> <p>In order to demonstrate the utility and convenience of Oxford Nanopore Technologies’ sequencing platform for performing detection and analysis of 5mC, we have sequenced the HG002 Genome in a Bottle Sample GM24385 with both traditional bisulfite sequencing and using nanopore sequencing. Both technologies, old and new, were applied to the same sample from a single DNA extraction.</p> <p><em>Bisulfite sequencing</em></p> <p>Bisulfite sequencing was performed by a commercial provider and processed with the commonly used <a href="https://www.bioinformatics.babraham.ac.uk/projects/bismark/">bismark</a> package to obtain the proportion of reads displaying methylation at CpG sites throughout the whole genome. </p> <p><em>Nanopore sequencing</em></p> <p>Nanopore sequencing was performed using the same sample of GM24385 material sent for bisulfite sequencing. Sequencing was performed on the MinION platform, across multiple flowcells, as part of ongoing platform development activities. The sequencing was not performed explicitly for the analysis presented here; we are making available all sequencing runs undertaken with this sample for the benefit of the community.</p> <p><em>Data Access</em></p> <p>Data is available as part of the Registry of Open Data on AWS: https://registry.opendata.aws/ont-open-data/. This dataset can be accessed through the S3 prefix:</p> <blockquote> <p>s3://ont-open-data/gm24385_mod_2021.09/</p> </blockquote> <p><em>Further Information</em></p> <ul> <li>https://labs.epi2me.io/gm24385-5mc/</li> <li>https://labs.epi2me.io/gm24385-5mc-remora</li> </ul>
Evaluation of a mNGS Workflow for Infection Diagnosis Using Oxford Nanopore Sequencing.
ClinicalTrials.gov study NCT04864873. IPD Sharing: NO. Countries: 1. Publications: 12.
Data from: A first look at the Oxford Nanopore MinION sequencer
Open the record for dataset details and reuse information.
Data for: Benchmarking Oxford Nanopore read assemblers for high-quality molluscan genomes
<p><span>Choosing the optimum assembly approach is essential to achieving a high-quality genome assembly suitable for comparative and evolutionary genomic investigations. Significant recent progress in long-read sequencing technologies such as PacBio and Oxford Nanopore Technologies (ONT) also brought about a large variety of assemblers. Although these have been extensively tested on model species such as <i>Homo sapiens </i>and <i>Drosophila melanogaster</i>, such benchmarking has not been done in Mollusca which lacks widely adopted model species. Molluscan genomes are notoriously rich in repeats and are often highly heterozygous, making their assembly challenging. Here, we benchmarked 10 assemblers based on ONT raw reads from two published molluscan genomes of differing properties, the gastropod <i>Chrysomallon squamiferum </i>(356.6Mb, 1.59% heterozygosity) and the bivalve <i>Mytilus coruscus</i> (1593Mb, 1.94% heterozygosity). By optimising the assembly pipeline, we greatly improved both genomes from previously published versions. Our results suggested that 40-50X of ONT reads are sufficient for high-quality genomes, with Flye being the recommended assembler for compact and less heterozygous genomes exemplified by <i>C. squamiferum</i>, while NextDenovo excelled for more repetitive and heterozygous molluscan genomes exemplified by <i>M. coruscus</i>. A phylogenomic analysis utilising the two updated genomes with other 32 published high-quality lophotrochozoan genomes resulted in maximum support across all nodes, and we show that improved genome quality also leads to more complete matrices for phylogenomic inferences. Our benchmarking will ensure the efficiency in future assemblies for molluscs and perhaps also other marine phyla with few genomes available.</span></p>
Can immature stages be ignored in studies of forest leaf litter arthropod diversity? A test using Oxford Nanopore DNA barcoding
<p>Datasets and results for the study</p>
Nanomotif: Identification and Exploitation of DNA Methylation Motifs in Metagenomes using Oxford Nanopore Sequencing
Open the record for dataset details and reuse information.
Supporting data for Comparison of Oxford Nanopore and Illumina sequencing for SARS-CoV-2 variant monitoring in wastewater
<p>This dataset contains the FASTQ files used in the paper <em>Comparison of Oxford Nanopore and Illumina sequencing for SARS-CoV-2 variant monitoring in wastewater</em>. </p> <p>The FASTQ files are classified in 2 different folders: MinION and MiSeq. Inside the MinION folder, there are two more folders, one for the R9.4.1 flow cell data and the other for the R10.4.1 flow cell data. </p> <p>Twist synthetic RNA mixtures are named as mix1, mix2, etc. </p> <p>Twist synthetic RNA corresponding to the Wuhan sequence is named as mix11_WH</p> <p>Wastewater samples are named as WWTP_1, WWTP_2, etc. </p>
Data for: Benchmarking Oxford Nanopore read assemblers for high-quality molluscan genomes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.