Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.9.0
Dataset results
26 results for “Oxford nanopore sequencing”
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing
<p>Curated dataset for the manuscript named "DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing".</p> <p>Five publicly available datasets sequenced on ONT MinION/GridION were used for the experiments (see below for original sources). These datasets contained raw signal data in single-FAST5 format (one file per each read), which were converted to BLOW5 format using slow5tools to enable convenient and efficient file manipulation. Then, 40,000 reads containing at least 4500 signal samples were extracted from each dataset. From each dataset, 20,000 reads are for training (<species>/train-<species>.blow5) and the rest for testing (<species>/test-<species>.blow5). Basecalled reads for the dataset used for testing are also available (test-<species>.fastq). Guppy version 6.1.3 under dna_r9.4.1_450bps_hac mode was used. The reference genomes are also given (<species>/<species>-ref.fasta)</p> <p>Original datasets are from the following sources:<br> SARS-CoV-2: https://community.artic.network/t/links-to-raw-fast5-fastq-data-for-artic-protocol/17<br> Zymo Metagenome: https://github.com/LomanLab/mockcommunity<br> Chlamydomonas: https://sra-download.ncbi.nlm.nih.gov/traces/era20/ERZ/003237/ERR3237140/Chlamydomonas_0.tar.gz<br> Saccharomyces cerevisiae: https://www.ncbi.nlm.nih.gov/bioproject/PRJNA510813</p>
Hieracium alpinun PAI33838 (2n = 2x = 18) Oxford Nanopore Technology sequences library
<p>Sample 1 000 000 reads (trimmed).</p>
Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing
<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>
Oxford Nanopore sequencing for comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity
<p><span>The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here, we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. </span></span></p>
Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns (repository for Genome Research paper, 2022)
<p>Simulated ONT and PacBio RNA-Seq data for "Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns" paper (Mikheenko et al., Genome Research, 2022). All details can be found in the Methods section of the paper.</p> <p><strong>PacBio.simulated_uniform_coverage.fasta.gz</strong> and <strong>ONT.simulated_uniform_coverage.fasta.gz files</strong> were used in Supplemental Note “Benchmarking of the read-to-isoform assignment algorithm”.</p> <p><strong>ONT.simulated_real_expression.fasta.gz</strong> file and all GTF files were used in the Section "Splice site correction improves transcript discovery precision". <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> was used as the annotation file for all tools. <strong>mouse.gencode.M26.spatial.15percent.expressed.gtf </strong>contains the set of all expressed isoforms. <strong>mouse.gencode.M26.spatial.15percent.expressed_kept.gtf</strong> contains those of the isoforms that are in presented in the annotation file ("known" transcripts), <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> contains expressed isoforms that were removed from the annotation ("novel" transcripts).</p>
Rapid and Inexpensive Whole-Genome Sequencing of SARS-CoV2 using 1200 bp Tiled Amplicons and Oxford Nanopore Rapid Barcoding
<p>Description of 1200bp amplicon primer sets and .bed and .tsv files for SARS-CoV-2 assembly using the ARTIC bioinformatics pipeline.</p>
Oxford Nanopore Direct RNA Sequencing datasets for detecting rRNA modifications in the Brassica oleracea mitoribosome
<p>Oxford Nanopore Direct RNA Sequencing (DRS) was applied for the detection of rRNA modifications in the <em>Brassica oleracea</em> mitoribosome. A comparison between native rRNA transcripts and in vitro transcribed (IVT) rRNA transcripts that were devoid of any modification indicated systematic base-calling errors and/or variations in current intensities that led to the prediction of the modified nucleotides (Begik et al., 2018).</p> <p>For Nanopore (DRS) library preparation, custom reverse transcription adapters (RTAs) containing Deeplexicon multiplexing barcodes (BC1, BC2 or BC3) were designed for sequence-specific ligation to the 3’-ends of <em>B. oleracea</em> mitochondrial rRNAs.</p> <p>Basecalling and demultiplexing of ONT direct RNA sequencing data were performed by Guppy and Deeplexicon, respectively.</p> <p>For the analysis and visualization of current intensities, a full description is available in the associated publication.</p> <p>The <strong>Eventalign_18S.R</strong> Rscript was used to make the nanopore <strong>18S </strong>signal analysis</p> <p>The <strong>Eventalign_26S.R</strong> Rscript was used to make the nanopore <strong>26S </strong>signal analysis</p> <p>The raw data for the scripts are in the following folders:</p> <p>- For 26S: 26S_Native.barcode01 and 26S_IVT.barcode03<br>- For 18S: 18S_Native.barcode01 and 18S_IVT.barcode02</p> <p>These folders contain the read alignment files output from minimap2 (in .bam format) and the eventalign files generated by the f5c software (in tsv format).</p> <p>The FASTA folder contains the mitochondrial 18S and 26S rRNA gene reference sequence in fasta format</p>
Matched Oxford Nanopore Technologies and Bisulfite Sequencing of the GM24385 Cell Line
<p>One of the most widespread genomic modifications is 5-methylcytosine (5mC), which most frequently occurs at <a href="https://en.wikipedia.org/wiki/CpG_site">CpG</a> dinucleotides. Compared to whole-genome bisulfite sequencing, the traditional method of 5mC detection, nanopore technology can offer many advantages such as simplicity of sample prep and subsequent analysis.</p> <p>In order to demonstrate the utility and convenience of Oxford Nanopore Technologies’ sequencing platform for performing detection and analysis of 5mC, we have sequenced the HG002 Genome in a Bottle Sample GM24385 with both traditional bisulfite sequencing and using nanopore sequencing. Both technologies, old and new, were applied to the same sample from a single DNA extraction.</p> <p><em>Bisulfite sequencing</em></p> <p>Bisulfite sequencing was performed by a commercial provider and processed with the commonly used <a href="https://www.bioinformatics.babraham.ac.uk/projects/bismark/">bismark</a> package to obtain the proportion of reads displaying methylation at CpG sites throughout the whole genome. </p> <p><em>Nanopore sequencing</em></p> <p>Nanopore sequencing was performed using the same sample of GM24385 material sent for bisulfite sequencing. Sequencing was performed on the MinION platform, across multiple flowcells, as part of ongoing platform development activities. The sequencing was not performed explicitly for the analysis presented here; we are making available all sequencing runs undertaken with this sample for the benefit of the community.</p> <p><em>Data Access</em></p> <p>Data is available as part of the Registry of Open Data on AWS: https://registry.opendata.aws/ont-open-data/. This dataset can be accessed through the S3 prefix:</p> <blockquote> <p>s3://ont-open-data/gm24385_mod_2021.09/</p> </blockquote> <p><em>Further Information</em></p> <ul> <li>https://labs.epi2me.io/gm24385-5mc/</li> <li>https://labs.epi2me.io/gm24385-5mc-remora</li> </ul>
Evaluation of a mNGS Workflow for Infection Diagnosis Using Oxford Nanopore Sequencing.
ClinicalTrials.gov study NCT04864873. IPD Sharing: NO. Countries: 1. Publications: 12.
Data from: A first look at the Oxford Nanopore MinION sequencer
Open the record for dataset details and reuse information.
Nanomotif: Identification and Exploitation of DNA Methylation Motifs in Metagenomes using Oxford Nanopore Sequencing
Open the record for dataset details and reuse information.
Supporting data for Comparison of Oxford Nanopore and Illumina sequencing for SARS-CoV-2 variant monitoring in wastewater
<p>This dataset contains the FASTQ files used in the paper <em>Comparison of Oxford Nanopore and Illumina sequencing for SARS-CoV-2 variant monitoring in wastewater</em>. </p> <p>The FASTQ files are classified in 2 different folders: MinION and MiSeq. Inside the MinION folder, there are two more folders, one for the R9.4.1 flow cell data and the other for the R10.4.1 flow cell data. </p> <p>Twist synthetic RNA mixtures are named as mix1, mix2, etc. </p> <p>Twist synthetic RNA corresponding to the Wuhan sequence is named as mix11_WH</p> <p>Wastewater samples are named as WWTP_1, WWTP_2, etc. </p>
Generation of full-length circRNA libraries for Oxford Nanopore long-read sequencing
GEO Series GSE197872. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
CircRNAs full-length sequencing in the human and mouse brain samples using Oxford Nanopore Technology
GEO Series GSE127059. Homo sapiens; Mus musculus. 8 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Profiling the Full-Length Transcriptome of Extracellular Vesicles With Oxford Nanopore Sequencing
GEO Series GSE225471. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Pore-C: Combination of chromatin capture assay and Oxford Nanopore Technology long read sequencing.
GEO Series GSE149117. Homo sapiens. 5 samples. Type: Other.
Detection of SINEUP RNA modification by Oxford Nanopore direct-RNA sequencing
GEO Series GSE224018. Homo sapiens. 3 samples. Type: Other.
Unraveling the whole genome DNA methylation profile of zebrafish kidney marrow through Oxford Nanopore sequencing
GEO Series GSE232842. Danio rerio. 12 samples. Type: Methylation profiling by high throughput sequencing.
scNanoATAC-seq: Long-read Single-cell ATAC-seq by Oxford Nanopore Technologies Sequencing
GEO Series GSE194024. Mus musculus; Homo sapiens. 17 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.