Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

181

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

181 results for “Nanopore sequencing”

Learn how ShareScore rates datasets ↗
zenodo40/100

De novo nanopore sequencing overrepresents RNA modification landscape

<p>RNA modifications are critical to the functional diversity and regulatory complexity of the transcriptome. With increasing frequency, direct nanopore RNA sequencing is applied to identify RNA modifications&nbsp;de novo. Here, we directly compare the MS2 phage genome RNA modification profiles determined using nanopore to orthogonal LC-MS/MS assays. The results reveal very different views of the modification landscape, suggesting caution when calling new RNA modifications using nanopore alone.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon" - PART I

<p>Raw Fast5 data for &quot;Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon&quot;. See Supplementary Table 2 for associating each sample to its barcode.</p> <p>- FC1_1 includes data for the HM mock community from BEI resources and skin microbiome of the chin in dogs.</p> <p>- FC1_2 includes data for the dorsal skin samples</p> <p>- FC2 includes data for the Zymobiomics mock community&nbsp;and Staphylococcus pseudintermedius isolate</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo36/100

Rapid and real-time identification of fungi up to the species level with long amplicon Nanopore sequencing from clinical samples

<p>Samples collected from fungal cultures, skin of dogs and ZymoBIOMICS<sup>TM </sup>mock community (which includes <em>Saccharomyces cerevisiae</em> and <em>Cryptococcus neoformans</em>). The amplicons length of the fungal cultures and&nbsp;ZymoBIOMICS<sup>TM </sup>mock community is 3,5 Kb and 6 Kb, while the <em>Malassezia spp</em> samples used as control is 3,5 Kb. The amplicons length of the four samples from the skin is 3,5 Kb.</p>

opencc-by-4.0Feb 2020View details →
zenodo36/100

SMAdd-seq: Probing chromatin accessibility with small molecule DNA intercalation and nanopore sequencing

<p>Studies of in vivo chromatin organization have relied on the accessibility of the underlying DNA to nucleases or methyltransferases, which is limited by their requirement for purified nuclei and enzymatic treatment. Here, we introduce a nanopore-based sequencing technique called Small-Molecule Adduct sequencing (SMAdd-seq), where we profile chromatin accessibility by treating nuclei or intact cells with a small molecule, angelicin. Angelicin reacts with thymine bases in linker DNA not bound to core nucleosomes after UV light exposure, thereby labeling accessible DNA regions. By applying SMAdd-seq in Saccharomyces cerevisiae, we demonstrate that angelicin-modified DNA can be detected by its distinct nanopore current signals. To systematically identify angelicin modifications and analyze chromatin structure, we developed a neural network model, NEural network for mapping MOdifications in nanopore long-reads (NEMO). NEMO accurately called expected nucleosome occupancy patterns near transcription start sites at both bulk and single-molecule levels. We observe heterogeneity in chromatin structure and identify clusters of single-molecule reads with varying configurations at specific yeast loci. Furthermore, SMAdd-seq performs equivalently on purified yeast nuclei and intact cells, indicating the promise of this method for in vivo chromatin labeling on long single molecules to measure native chromatin dynamics and heterogeneity.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing

<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Oxford Nanopore sequencing for comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity

<p><span>The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here, we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. </span></span></p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

NASTRA: Accurate analysis of short tandem repeat markers by nanopore sequencing with repeat-structure-aware algorithm

<p><span>Forensic short-tandem repeats (STR) genetic markers are multi-allelic and widely utilized for individual identification, kinship testing, and cell-line authentication. Nanopore sequencing, known for its portability, is emerging as a promising approach for STR typing, facilitating real-time and in-field testing. However, its efficacy is often hampered by sequencing noise. Previous methods rely on alignment-based genotyping, necessitating known alleles, which limits their applicability to unknown alleles. Here, we introduced NASTRA, an innovative allele reference-free tool for precise germline analysis of STR genetic markers. NASTRA incorporates a recursive algorithm to infer repeat structures of allele sequences using only known repeat motifs. Our tests, conducted on 80 individual samples and 8 DNA standards, have demonstrated NASTRA's exceptional 100% accuracy in genotyping nearly all diploid STRs across various multiplex kits and flow cells. It surpasses alignment-based methods in accuracy and speed. In a paternity testing case study, NASTRA accurately identified three relationships among six individuals within an 18-minute sequencing duration. These results underscore NASTRA's ability to perform STR analysis on both NGS and nanopore sequencing platforms, significantly enhancing the utility of nanopore sequencing in relevant applications.</span></p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns (repository for Genome Research paper, 2022)

<p>Simulated ONT and PacBio RNA-Seq data for &quot;Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns&quot; paper (Mikheenko et al., Genome Research, 2022). All&nbsp;details can be found in the Methods section of the paper.</p> <p><strong>PacBio.simulated_uniform_coverage.fasta.gz</strong>&nbsp;and <strong>ONT.simulated_uniform_coverage.fasta.gz&nbsp;files</strong> were used in&nbsp;Supplemental Note &ldquo;Benchmarking of the read-to-isoform assignment algorithm&rdquo;.</p> <p><strong>ONT.simulated_real_expression.fasta.gz</strong>&nbsp;file and all GTF files were used in the Section &quot;Splice site correction improves transcript discovery precision&quot;.&nbsp;<strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong>&nbsp;was used as the annotation file for all tools.&nbsp;<strong>mouse.gencode.M26.spatial.15percent.expressed.gtf </strong>contains the set of all expressed isoforms.&nbsp;<strong>mouse.gencode.M26.spatial.15percent.expressed_kept.gtf</strong> contains those&nbsp;of the&nbsp;isoforms that are in presented in the annotation file (&quot;known&quot; transcripts),&nbsp;<strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> contains expressed isoforms that were removed from the annotation&nbsp;(&quot;novel&quot; transcripts).</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

Nanopore sequencing data analysis using Microsoft Azure cloud computing service

<p>Genetic information provides insights into the exome, genome, epigenetics and structural organisation of the organism. Given the enormous amount of genetic information, scientists are able to perform mammoth tasks to improve the standard of health care such as determining genetic influences on outcome of allogeneic transplantation. Cloud-based computing has increasingly become a key choice for many scientists, engineers and institutions as it offers on-demand network access and users can conveniently rent rather than buy all required computing resources. With the positive advancements of cloud computing and nanopore sequencing data output, we were motivated to develop an automated and scalable analysis pipeline utilizing cloud infrastructure in Microsoft Azure to accelerate HLA genotyping service and improve the efficiency of the workflow at lower cost. In this study, we describe (i) the selection process for suitable virtual machine sizes for computing resources to balance between the best performance versus cost-effectiveness; (ii) the building of Docker containers to include all tools in the cloud computational environment; (iii) the comparison of HLA genotype concordance between the in-house manual method and the automated cloud-based pipeline to assess data accuracy. In conclusion, the Microsoft Azure cloud-based data analysis pipeline was shown to meet all the key imperatives for performance, cost, usability, simplicity and accuracy. Importantly, the pipeline allows for the ongoing maintenance and testing of version changes before implementation. This pipeline is suitable for data analysis from MinION sequencing platforms and could be adopted for other data analysis application processes.</p>

opencc-zeroOct 2022View details →
zenodo36/100

Temperature modulates dominance of a superinfecting Arctic virus in its unicellular algal host - Nanopore sequencing reads

<p>Nanopore sequencing reads of two Micromonas polaris viruses, MpoV-45T and MpoV-46T using R9 chemistry.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Experimental data for "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads"

<p>The experimental dataset used in "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads."</p> <p>A set of 91,766 150-nt oligos were synthesised with GenScript (oligos.fasta). Each oligo consists of a pseudo-random 110-nt payload flanked by 20-nt primers at each end. The strands are split in three roughly equal groups (two groups of 30,589 and one group of 30,588). Each group has a dedicated primer pair for targeted PCR amplification (the primer pairs used for amplification are provided in primers_synthesis.fasta). The pseudo-random payload was designed to avoid primer-payload collisions.</p> <p>For each file, a sample from the synthesised pool was PCR amplified using the corresponding primer pair and sequenced using Oxford Nanopore Technologies MinION sequencing device following the standard library preparation protocol for amplicon DNA. The raw reads were basecalled using guppy, either in fast- ("acc-false") or high-accuracy ("acc-true") regime. The basecaller generated two groups of reads&mdash;"passQ-true" for the reads that passed the quality-score threshold of 8 and "passQ-false" for those that did not. For each group of reads, a BLAST-based fuzzy search for primer sequences was performed and, based on the resulting alignments, the segments containing the correct primer pairs and located at a distance of 150+-15nt were extracted (separately for forward and reverse-complemented reads). The segments are then assigned to the closest synthesized strand based on Levenshtein distance. The resulting clusters are used to estimate the parameters of the end-to-end DNA storage channel model and to test the proposed error-correction scheme.</p> <p>The archive clustered_read_segments.tar.gz contains 12 sub-archives, for each file (0,1,2), accuracy ("acc-true" or "acc-false"), and Q-score ("passQ-true" or "passQ-false"). Within each sub-archive, there are two folders (one for forward read segments and one for backward read segments), and each folder contains two files: one for the reference synthesised (or "transmitted") sequences that correspond to the file in question ("TX__" &mdash; e.g., "TX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt") and another file for the sequenced (or "received") segment clusters ("RX__" &mdash; e.g., "RX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt"). The received clusters in the "RX__" file are ordered in correspondence with the synthesised sequences in the "TX__" file, and a line "===============================" is used as a separator.</p>

opencc-by-4.0May 2024View details →
dryad36/100

Nanopore sequencing assay to detect and diagnose tuberculous meningitis via cerebrospinal fluid

<p>This study aimed to evaluate the efficiency of nanopore sequencing for the early diagnosis of tuberculous meningitis (TBM) using cerebrospinal fluid and compared it with acid-fast bacilli (AFB) smear, mycobacterial growth indicator tube (MGIT) culture, and Xpert MTB/Rifampicin (RIF). We enrolled 64 adult patients with presumptive TBM admitted to our hospital from August 2021 to August 2023. We calculated the sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) of AFB smear, culture, Xpert MTB/RIF, and nanopore sequencing to evaluate their diagnostic efficacy compared with a composite reference standard for TBM. Among these 64 patients, all tested negative for TBM by AFB smear. The sensitivity, specificity, PPV, and NPV were 11.11%, 100%, 100%, and 32.2% for culture, 13.33%, 100%, 100%, and 2.76% for Xpert MTB/RIF, and 77.78%, 100%, 100% and 65.52% for nanopore sequencing, respectively. The diagnostic accuracy of the nanopore sequencing test was significantly higher than that of conventional testing methods used to detect TBM.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Control panel created from 30-40 Nanopore or PacBio HiFi sequencing data from the Human Pangenome Reference Consortium

<p>This is control panel for <a href="https://github.com/friend1ws/nanomonsv">nanomonsv</a> software, which is expected to exclude many false positives as well as improve computational cost. This is made by aligning 30-40 Nanopore or PacBio HiFi sequencing data from Human Pangenome Reference Consortium (HPRC) to the GRCh38 or CHM13 reference genomes with <a href="https://github.com/lh3/minimap2">minimap2</a> version 2.24.</p> <p><strong>When you use these control panels and publish, do not forget to credit to <a href="https://humanpangenome.org/data-use-protocol/">HPRC</a>!</strong></p> <div> <div> <div> <p>Reference genomes:</p> <ul> <li>GRCh38: <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz">Download GRCh38</a></li> <li>CHM13: <a href="https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0_maskedY_rCRS.fa.gz">Download CHM13</a> <div> <div> <div> <div>&nbsp;</div> </div> </div> </div> </li> </ul> </div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Nanopore sequence analysis - Galaxy Training Material

<p>Twelve MDR plasmids harboring samples were prepared according to the MinION library construction protocols, followed by library sequencing. After 8 hours of sequencing run, a total of 287 725 reads ranging from dozens to tens of thousands of bases in length were obtained, covering a total of 493 Mbp. The raw data were subjected to several stages of processing, including basecalling, de-multiplexing, fasta sequence extraction. For this tutorial one out of the twelve samples is chosen as example.</p> <p>This dataset is extracted of a&nbsp;project studying the Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data (<a href="https://doi.org/10.1093/gigascience/gix132">https://doi.org/10.1093/gigascience/gix132</a>)</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

NanoBaseLib: A Multi-Task Benchmark Dataset for Nanopore Sequencing

<p>NanoBaseLib is a multi-task benchmark dataset for Nanopore Sequencing. We compile and preprocess publicly available datasets using a unified pipeline to ensure consistency and quality across all tasks. The dataset is benchmarked for four key Nanopore sequencing tasks: base calling, polyA detection, segmentation and event alignment, and RNA modification detection. &nbsp;NanoBaseLib is available at <a href="https://nanobaselib.github.io/">https://nanobaselib.github.io</a>.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

NPIP: A Comprehensive Analysis Pipeline for Rapid Pathogen Detection in Clinical Samples Based on Nanopore Sequencing

<p>Background: Rapid and accurate pathogen detection is important for effective control of infectious diseases. However, traditional pathogen culture methods have a very long detection time, as well as high rates of false-positive and false-negative results. Third generation sequencing (TGS) technology brings the new possibility of being used as a pathogen detection method. However, the practicability of a pathogen detection report based on TGS is still lacking. There is also a lack of professional and accurate report interpretation.</p> <p>Results: Here, we report on the development of a pathogen detection and analysis tool (NPIP) based on third generation nanopore sequencing technology. We also prove the practicability of nanopore sequencing and NPIP analysis tools in emergency and clinical pathogen detection by demonstrating its use in a practical case.</p> <p>Conclusions: This platform provides an effective, convenient, and fast analysis tool for clinicians and public health personnel to more successfully apply TGS in pathogen detection.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Direct nanopore sequencing of human cytomegalovirus ge-nomes from high-titre clinical samples

<p>This archive contains the nanopore generated HCMV genomes from the urine and lung clinical samples. The illumina derived genome sequences have been uploaded to GenBank as they are of higher quality, the nanopore sequences are uploaded here to avoid duplication on GenBank. Raw FASTQ reads from the virus are uploaded to NCBI SRA.&nbsp;The full paper abstract is given below.&nbsp;Nanopore sequencing is becoming increasingly commonplace in clinical settings, particularly for diagnostics and outbreak investigations. Its portability, low cost and ability to operate in near real-time has propelled it to the forefront of the recent SARS-CoV-2 pandemic. Although high sequencing error rates initially hampered its wider implementation, improvements have continually been made with each iteration of the nanopore flow cells and base calling software. Here, we assess the feasibility of using nanopore sequencing to determine the complete genome of human cytomegalovirus (HCMV) present in high-titre clinical samples without viral DNA&nbsp;enrichment, PCR amplification or prior knowledge of the sequences. We utilised a hybrid bioinformatic approach that involved assembling the reads<em>&nbsp;de novo</em>, improving the&nbsp;<em>de novo</em>&nbsp;consensus through alignment of reads to the best-matching genome from a collated set of published genomes, and polishing the improved consensus. The final consensus genome sequences from a urine and a lung sample, the latter with an HCMV to human DNA load approximately 50 times lower than the former, achieved 99.97 and 99.93% identity, respectively, to the bench-mark consensuses obtained independently by Illumina sequencing. Thus, we demonstrate that nanopore sequencing is capable of determining HCMV genomes directly from high-titre clinical samples with high accuracy.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
dryad36/100

Nanopore sequencing genomic DNA from Stentor pyriformis and its endosymbiont, Chlorella variabilis

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Using cerebrospinal fluid nanopore sequencing assay to diagnose tuberculous meningitis: a retrospective cohort study in China

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

Nanopore sequencing data analysis using Microsoft Azure cloud computing service

Open the record for dataset details and reuse information.

publicOct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record