Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
181
datasets available to search
ShareScore release 0.9.0
Dataset results
181 results for “Nanopore sequencing”
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
Tracking Down Chimeric Assemblies In The TrackIt DNA Ladder Using Nanopore Sequencing
<h2>Dataset Description</h2><p>These files represent two different LSK114 sequencing runs on a TrackIt 1kb Plus DNA Ladder sample, and associated data analysis.</p><ul><li>July 20 2023 Flongle Run (191 Mb; 465k reads)<ul><li>pod5_files_2023-Jul-20_DAE_DNA_Ladder.tar.gz<br>- raw POD5 format files</li><li>called_2023-Jul-20_DAE_DNA_Ladder_duplex.bam<br>- duplex called reads, called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Jul-20_DAE_DNA_Ladder.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Jul-20_DAE_DNA_Ladder_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Jul-20_DAE_DNA_Ladder.txt<br>- Length / QC summary statistics</li></ul></li><li>October 12 2023 P2 Solo Run (1.95 Gb, 1.11M reads)<ul><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_fail.tar.gz<br>- raw POD5 format files (all failed reads)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_000-059.tar.gz<br>- raw POD5 format files (passed reads, bundle #000-059)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_060-119.tar.gz<br>- raw POD5 format files (passed reads, bundle #060-119)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_120-179.tar.gz<br>- raw POD5 format files (passed reads, bundle #120-179)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_180-222.tar.gz<br>- raw POD5 format files (passed reads, bundle #180-222)</li><li>called_2023-Oct-12_DNA-Ladder-1kbplus_duplex.bam<br>- duplex called reads [October 12, 2023], called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Oct-12_DNA-Ladder-1kbplus.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Oct-12_DNA-Ladder-1kbplus_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Oct-12_DNA-Ladder-1kbplus.txt<br>- Length / QC summary statistics</li><li>ladder_seqs.fa<br>- assembled DNA ladder sequences, based on simplex reads</li></ul></li></ul><h3>Methods</h3><h3>Sample preparation</h3><p>Preparation of DNA for sequencing was carried out following the ONT Ligation Sequencing DNA V14 (SQK-LSK114) protocol, with modifications to exclude DNA repair, and keeping the sample in the same 1.5ml tube to reduce sample loss.</p><h4>Tris-buffered Saline (TBS) buffer preparation</h4><ol><li>1M stock of NaCl was made by adding 2.922g of NaCl into a 50 ml Falcon tube, then made up to 50 ml with MilliPore water</li><li>A 50 mM TBS stock was created by adding 750 μl 1M NaCl solution to a 15 ml Falcon tube, then made up to 15 ml using Qiagen Elution Buffer (EB, i.e. 10 mM Tris-HCl at pH 8.0)</li><li>The pH was confirmed to be 7.9-8.1 using a pH indicator strip (e.g. MColorpHast 6.5 - 10.0; MER1095430001)</li></ol><h4>End prep</h4><ol><li>1 μg DNA ladder (i.e. 10 μl of 0.1 μg / μl DNA ladder) was transferred into a 1.5ml Eppendorf DNA LoBind tube</li><li>The volume was topped up to 43.5 μl with TBS (i.e. 33.5 μl TBS)</li><li>3.5 μl Ultra II End-prep Reaction Buffer and 3 μl Ultra II End-prep Enzyme Mix was added</li><li>After mixing by gentle pipetting, the mixture was incubated at RT for 5 minutes, then 65 \degrees for 5 minutes</li></ol><h4>Bead cleanup</h4><ol><li>The mixture was combined with 60 μl Ampure XP beads, and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 150 μl of an 80% ethanol solution</li><li>The sample was dried briefly for 30s, then eluted in 60 μl TBS</li></ol><h4>Adapter ligation and final bead cleanup</h4><ol><li>To the sample tube was added 25 μl ONT Ligation buffer (LNB), 5μl NEBNext Quick T4 DNA Ligase (reduced from the protocol-suggested 10μl because that was all that was left in the tube), and 5μl ONT Ligation Adapter (LA)</li><li>The tube was mixed by gentle pipetting, spun down for 1-3s on a mini centrifuge, then incubated for 10 minutes at RT</li><li>The mixture was combined with 40 μl Ampure XP beads (100μl Ampure XP beads were used for the Flongle sample), and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 250 μl of ONT Long Fragment Buffer (LFB) for the P2 Solo run, and 250μl ONT Short Fragment Buffer <br>(SFB) for the Flongle run</li><li>The sample was dried briefly for 30s, then eluted for 10 minutes at 37 \degrees in 15 μl ONT Elution buffer (EB)</li></ol><h4>Addition of sequencing library buffers</h4><ol><li>A flow cell was prepared by flushing with ONT Flow Cell Flush (FCF) mixed with ONT Flow Cell Tether (FCT). For the P2 Solo, I used 500 μl of a 1170μl FCF solution that had 30μl FCT added to it; for the Flongle I used 60μl of a 117μl FCF solution that had 3 μl FCT added to it</li><li>1 μl of the eluted library was quantified on a Quantus Fluorometer, and approximately 50 fmol (assuming 1kb average length) was transferred to a new 1.5μl tube</li><li>For the P2 Solo run, the volume was topped up to 32 μl TBS; for the Flongle run, the volume was topped up to 12 μl TBS</li><li>To the sample tube was added ONT Sequencing Buffer (SB; P2 Solo - 100μl; Flongle - 30μl) and ONT Library Beads (LIB; P2 Solo - 68μl; Flongle - 20μl)</li><li>The flow cell was re-flushed with additional FCF/FCT mixture (500 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The sequencing library was then added to the flow cell (200 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The prepared flow cell was left for 10 minutes to allow the library to settle before starting sequencing</li></ol><h4>DNA Sequencing and basecalling</h4><ol><li>Sequencing was carried out using MinKNOW v23.04.6, sequencing in fast mode at 400 bases per second with a 20bp minimum sequence length and <br>5 kHz sampling rate, with reads output as POD5 files</li><li>The Flongle flow cell was run for a full standard run length (24h), whereas the PromethION flow cell was run for 1.5 hours (after which the <br>counts of 15kb reads exceeded 200)</li><li>Sequenced reads were recalled in standard (simplex) mode using Dorado v0.4.0 and the 2023-09-22 bacterial methylation model [res_dna_r10.4.1_e8.2_400bps_sup@2023-09-22_bacterial-methylation]</li></ol><h3>Bioinformatics Analysis of Ladder Sequences </h3><h4>Sequence assembly</h4><p>Assembly process for bands that are 3k in length and greater (done on LFB-depleted P2 Solo sequences): </p><ol><li>Filter >q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not sufficient for the 15kb band; all reads were needed]</li><li>Chop the reads up with a 1000bp overlap (e.g. 3000bp for the 5k band). This works around a Canu expectation that any read overlaps should be less than X% of the read.</li><li>Assemble the reads with Canu v2.2 [#REF], treating them as "pacbio" reads (for correction and homopolymer compression), with the GenomeSize parameter set to the expected band length (e.g. GenomeSize=5000).</li><li>Extract the first reported assembled contig.</li><li>Map the contig to the nanopore adapter sequences, and trim to exclude any matching sequence.</li></ol><p>[Canu has a default genome size and read length cutoff of 1kb, and performs poorly on sequences shorter than this] <br><br>Assembly process for bands under 3k in length (done on LFB-depleted P2 Solo sequences):</p><ol><li>Filter >q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not in sufficient abundance for the 100bp band; all reads were needed]</li><li>Assemble using a<a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kassembler.pl"> kmer-based de-bruijn assembler</a>, trimming off low-count kmers</li><li>Extract the first reported trimmed assembled chain</li><li>Map the assembled chain to the nanopore adapter sequences, and trim to exclude any matching sequence</li><li>Use web BLASTn [#REF] to help trim any additional trailing non-matching sequence</li></ol><h4>Mapping</h4><ol><li>Use a <a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kmapper.pl">kmer-based lightweight mapper</a> to map reads to assembled bands</li><li>Created LAST mismatch matrix using `last-train` on the 5k reads together, using the full assembled ladder sequences as a reference:<br>last-train -Q 1 ladder_seqs.fa 5k_reads.fq.gz</li><li>Mapped all reads to the assembled ladder sequences (only the reference corresponding to the most likely band source), retaining (for each read) the mapping that had the longest combined proportion of read and reference sequence mapped:<br>lastal -p bacterial.mat -P 10 ladder_seqs.fa reads_2023-Oct-12_DNA-Ladder-1kbplus_called_all.fq.gz | \ <br> ~/scripts/maf2csv.pl | \ <br> awk -F ',' '{print $0","($8/100 * $13/100)}' | \ <br> sort -t ',' -k 16rg,16 | sort -t ',' -k 1,1 -u | sort -t ',' -k 1r,1 | \ <br> perl -pe 's/,[^,]*$/\n/' > LAST_reads_vs_ladder_longestMatch.csv.gz</li></ol>
Control panel created from 30 Nanopore sequencing data from the Human Pangenome Reference Consortium
<p>This is control panel for <a href="https://github.com/friend1ws/nanomonsv">nanomonsv</a> software, which is expected to exclude many false positives as well as improve computational cost. This is made by aligning 30 Nanopore sequencing data from Human Pangenome Reference Consortium to the GRCh38 reference genome (obtained from <a href="https://console.cloud.google.com/storage/browser/genomics-public-data/resources/broad/hg38/v0;tab=objects">here</a>) with <a href="https://github.com/lh3/minimap2">minimap2</a> version 2.24. <strong>When you use these control panels and publish, do not forget to credit to <a href="https://humanpangenome.org/data-use-protocol/">HPRC</a>!</strong></p>
Fly through a 2.3Mb Nanopore Sequence
<p>A spiral representation of the longest sequence found by <a href="https://doi.org/10.1101/312256">Alex Payne et al.</a>, which at the time of upload was the longest observed public single-pore nanopore sequence. This video shows about 30kb of the sequence at once, and zooms through the sequence at about 20kb per second in the animated GIF, or about 48kb per second in the AVI file.</p>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing - SGNEx data
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
CyclomicsSeq: Accurate detection of circulating tumor DNA using nanopore consensus sequencing
<p>CyclomicsSeq is a protocol designed to produce and sequence long DNA concatemers with a linear repetition to acquire high accuracy consensus reads. In this dataset, we used CyclomicsSeq for sequencing TP53 in cell-free DNA of healthy individuals and of head and neck cancer patients and for sequencing synthetic TP53 DNA sequences that mimic the length of cell-free DNA. This dataset contains data (mainly base calls of the backbone and the insert) of 32 nanopore sequencing runs. <br> </p>
Alignment files for coverage benchmarks: Illumina and Nanopore sequencing datasets
<ul> <li><strong>cpara-illumina-noseq.bam</strong> and <strong>cpara-ont-noseq.bam</strong>: BAM files produced aligning the raw reads produced respectively by Illumina NextSeq and ONT Nanopore sequencing of an isolate of <em>C. parapsilosis</em> to evaluate the coverage calculations using real datasets.*</li> <li><strong>HG00258.bam</strong>: Exome sequencing from the 1000 Genomes Project (Clarke et al 2016 <a href="https://doi.org/10.1093/nar/gkw829">https://doi.org/10.1093/nar/gkw829</a>).</li> <li><strong>panel_01.bam</strong>: targeted sequencing of a Human gene panel of 16 genes.*</li> </ul> <p>* Sequences and qualities have been removed</p>
Genome-wide mapping of individual replication fork velocities using nanopore sequencing
<p>These data are related to the <strong>"Genome-wide mapping of individual replication fork velocities using nanopore sequencing" </strong>manuscript by Theulot, Lacroix et al. (https://doi.org/10.1038/s41467-022-31012-0) and to the GitHub repository NanoForkSpeed (https://github.com/LacroixLaurent/NanoForkSpeed).</p> <p>BT1_run4.tar.gz contains the nanopore reads from a typical experiment as fast5 files.</p> <p>BT1_run4_mega.zip contains the result of the BrdU basecalling as described on the GitHub page in a modified bam file format as described on the GitHub page NanoForkSpeed/BrdU_Basecalling..</p> <p>BT1_run4_Megalodon_00_smdata.rds contains the result of the parsing function for the modified bam file as described on the GitHub page NanoForkSpeed/BrdU_Basecalling.</p> <p>BT1_run4_Megalodon_00_NFS_data.rds contains the result of the fork detection procedure on the BT1_run4_Megalodon_00_smdata.rds file as described on the GitHub page NanoForkSpeed/Forks_Detection.</p> <p>BT1_run4_merged_NFS_data.rds contains the result of the NFS_merge function as described on the GitHub page NanoForkSpeed/Forks_Detection.</p> <p> </p>
Point-of-care monitoring of head and neck cancer treatment response and recurrence development using nanopore-based ctDNA consensus sequencing
<p>Circulating tumor DNA (ctDNA) in blood may become a generic biomarker for non-invasive cancer diagnosis and monitoring. However, detection of ctDNA is challenged by the presence of many circulating DNA molecules from healthy cells. We found that single ctDNA molecules can be sequenced with high accuracy by a three-step process consisting of capturing, copying and concatenation of the original double-stranded ctDNA molecules. This innovative approach - called CyclomicsSeq - is unparalleled by any other method in terms of cost-efficiency and speed, allowing point-of-care cancer diagnostics.</p> <p>Within this CPOC, subsidized by the Oncode institute, we have applied our CyclomicsSeq ctDNA test in patients with advanced head and neck cancer squamous cell carcinoma (HNSCC). Head and neck cancer (HNSCC) accounts for 380,000 cancer-related deaths worldwide. For these patients, determining whether a patient responds to the primary chemoradiation treatment is challenging, and non-responders are sometimes identified when other treatment options are no longer possible. By measuring the ctDNA levels in the blood of these patients prior to and during treatment, we aim to identify non-responders at an earlier stage.</p> <p>This dataset contains base calls of TP53 of 47 nanopore sequencing runs. We included 10 patients and 7 controls. For the patients, we have samples of multiple time points (0 = prior to treatment, 1 = 1 week after treatment initiation, etc).</p>
BeerDEcoded Nanopore Sequencing workshop of la trappe Dez. 2019
<p>The BeerDEcoded workshops are a project of the Street Science Community. https://streetscience.community/projects/beerdecoded/<br> We are extracting DNA from different Beers and sequence them using the Nanopore sequencing technique. The data contains the reads extracted from the beer of the brand la trappe and was performed on 08.12.2019. <br> The sequencing was performed on a MinION using a Flow Cell and the MinKNOW software. The Software collects sequencing data in real-time and has a base-calling integrated to convert the fast5 files into fastq files.</p>
BeerDEcoded Nanopore Sequencing Run of Chimay Now. 2019
<p>The BeerDEcoded workshops are a project of the Street Science Community. https://streetscience.community/projects/beerdecoded/<br> We are extracting DNA from different Beers and sequence them using the Nanopore sequencing technique. The data contains the reads extracted from the beer of the brand Chimay and was performed on 26.11.2019. <br> The Sequencing was performed on a MinION using a Flow Cell and the MinKNOW software. The Software collects sequencing data in real-time and has a base-calling integrated to convert the fast5 files into fastq files.</p>
BeerDEcoded Nanopore Sequencing Run of Chimay Now. 2019 Flongle
<p>The BeerDEcoded workshops are a project of the Street Science Community. https://streetscience.community/projects/beerdecoded/<br> We are extracting DNA from different Beers and sequence them using the Nanopore sequencing technique. The data contains the reads extracted from the beer of the brand Chimay and was performed on 26.11.2019. <br> The Sequencing was performed on a MinION using a flongle Flow Cell and the MinKNOW software. The Software collects sequencing data in real-time and has a base-calling integrated to convert the fast5 files into fastq files.</p>
DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing
<p>Curated dataset for the manuscript named "DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing".</p> <p>Five publicly available datasets sequenced on ONT MinION/GridION were used for the experiments (see below for original sources). These datasets contained raw signal data in single-FAST5 format (one file per each read), which were converted to BLOW5 format using slow5tools to enable convenient and efficient file manipulation. Then, 40,000 reads containing at least 4500 signal samples were extracted from each dataset. From each dataset, 20,000 reads are for training (<species>/train-<species>.blow5) and the rest for testing (<species>/test-<species>.blow5). Basecalled reads for the dataset used for testing are also available (test-<species>.fastq). Guppy version 6.1.3 under dna_r9.4.1_450bps_hac mode was used. The reference genomes are also given (<species>/<species>-ref.fasta)</p> <p>Original datasets are from the following sources:<br> SARS-CoV-2: https://community.artic.network/t/links-to-raw-fast5-fastq-data-for-artic-protocol/17<br> Zymo Metagenome: https://github.com/LomanLab/mockcommunity<br> Chlamydomonas: https://sra-download.ncbi.nlm.nih.gov/traces/era20/ERZ/003237/ERR3237140/Chlamydomonas_0.tar.gz<br> Saccharomyces cerevisiae: https://www.ncbi.nlm.nih.gov/bioproject/PRJNA510813</p>
Nanotiming: single-molecule based, telomere-to-telomere DNA replication timing profiling by nanopore sequencing
<p>Dataset for the manuscript "Nanotiming: telomere-to-telomere DNA replication timing profiling by nanopore sequencing" by Theulot et al ,2024 (<span>https://doi.org/10.1038/s41467-024-55520-3</span>) related to the github repository (https://github.com/LacroixLaurent/NanoTiming)</p> <ul> <li>WT_rep3.tar.gz contains fast5 file from an experiment where yeast BT1 strain was grown for one doubling time with 5µM BrdU then DNA was sequenced on R9.4.1 ONT flowcell</li> <li>mod_mapping.bam contains the bam file resulting from the BrdU base calling with megalodon (v2.2.9) using our BT1 reference genome and our BrdU aware model for base-calling</li> <li>WT_rep3_nanoT.bed.gz contains the reads coordinates from the mod_mappings file</li> <li>WT_rep3_nanoT_alldata.rds contains the BrdU profiles for each reads of the mod_mappings file, with the BrdU signal binned in 1kb non overlaping windows</li> <li>WT_rep3_nanoT.rds contains the genomic BrdU signal profiles by 1kb non overlaping windows</li> <li>TeloLengthDataNanoT.rds contains all the telomeric sequences extracted from the experiments reported in the Figure 4 and S19 to S23 of the manuscript with the associated filtering information and nanotiming signal.</li> </ul> <p> </p>
Nanopore deep sequencing as a tool to characterize and quantify aberrant splicing caused by variants in inherited retinal dystrophy genes
Open the record for dataset details and reuse information.
Hieracium alpinun PAI33838 (2n = 2x = 18) Oxford Nanopore Technology sequences library
<p>Sample 1 000 000 reads (trimmed).</p>
Nanopore sequencing of plasmid cleavage fragments produced with type III CRISPR-associated nucleases NucC, Can1 and Can2
<p>Included datasets were generated in the study "<strong>Sequence-specific capture and concentration of viral RNA </strong><strong>by type III CRISPR system enhances diagnostic"</strong> by Nemudraia et al., 2022</p> <p> </p> <p>For questions contact: Artem Nemudryi (artem.nemudryi@gmail.com) or Blake Wiedenheft (bwiedenheft.com)</p>
Simultaneous profiling of histone modifications and DNA methylation via nanopore sequencing
<p>Datasets that contain a minimum of nanopore reads sufficient for hidden Markov model training and for evaluating the performance of our computational tool - nanoHiMe at simultaneously calling CpG and/or adenine methylation on individual nanopore reads.<em> Ecoli</em>_PCR_amplicons_100k.tgz, <em>Ecoli</em>_PCR_MSssI_100k.tar.gz and <em>Ecoli</em>_PCR_pA-Hia5_100k.tar.gz are used for training new parameters of the emission distributions of individual <em>k</em>-mers from DNA template without modification, with fully methylated CpGs, and with partially methylated adenines, respectively. nanoHiMe_H3K27me3.fast5.tgz are the nanopore sequencing reads from H3K27me3 nanoHiMe-seq experiments in GM12878 cells and used for evaluating the performance of nanoHiMe at jointly calling CpG and adenine methylation.</p>
Methylation-free E.coli nanopore sequencing (ONT R9.4.1) data set
<p>The data set consists of fast5 files divided into 5 zip files (fast5_[1-5].zip), a genome record (Ecoli_K12_MG1655.fasta), an Illumina assembly genome (illumina_contigs.fasta) and a fastq file from Guppy 5 (guppy_basecalled.fastq.gz). We sequenced the Ecoli non-methylated genomic DNA (D5016, Zymo Research) with an ONT MinION device. The sequencing libraries were prepared by fragmenting the genomic DNA using Covaris g-TUBE and a Ligation sequencing kit (SQK-LSK109, Oxford Nanopore) with Flow Cell chemistry R9.4.1. We also performed short-read Illumina sequencing on the same sample using the TruSeq PCR-free library preparation on the MiSeq sequencing platform (Illumina, USA), and constructed a draft assembly from the Illumina sequencing results using SPAdes v3.6.0. We also upload a reference genome directly obtained from the E.coli sample producer website. </p> <p>In addition, the data set contains two fastq files that produced by the Lokatt basecaller (lokatt_basecalled.fasta.gz) and local-trained Bonito basecaller (bonito_local_basecalled.fastq.gz ), respectively, which are used for benchmarking in the Lokatt basecaller paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.