Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
233
datasets available to search
ShareScore release 0.9.0
Dataset results
233 results for “long-read”
Dataset for "Nanopore-led long-read genome assembly of the Australian yabby, Cherax destructor"
<p>Intermediate_Assemblies.tar.gz: Intermediate genome assemblies e.g. raw wtdbg assembly (CD.raw.fa), polished wtdbg assembly (CD.cns.fa), 1st pilon polished assembly (CDF2_pilon1.fasta), 2nd pilon polished assembly (CDF2_pilon2.fasta) and RNA-scaffolded assembly (CDF2_pilon2_prna.fasta). Folders with run_"assembly name" are BUSCO output for each of the assembly.</p> <p>BRAKER2.tar.gz: BRAKER2 genome annotation output containing the initial set of predicted protein-coding genes as well as training intermediate files.</p> <p>BUSCOv3.tar.gz: BUSCO assessment of publicly available Decapod crustacean genome assemblies</p> <p>Cdes.filtered.codingseq: Filtered set of protein-coding sequences</p> <p>Cdes.filtered.faa: Translation of the filtered protein-coding sequences</p> <p>CDF2.NCBI.fasta.masked.gz: Repeat-masked (softmasked) Cherax destructor genome</p> <p>Cqua_transcriptome.tar.gz: rnaSPAdes output (combined fasta) of all Cherax quadricarinatus transcriptomes and its reduced dataset generated by EvidentialGene. </p> <p>Quast.tar.gz: Quast output of all Decapod crustacean genome assemblies assessed in this study</p> <p>Repeat_Annotation.tar.gz: Repeat annotation (.gff3) based on Cherax destructor-specific de novo repeat library and its summary (.tbl)</p> <p>RepeatLibrary.tar.gz: Cherax destructor-specific de novo repeat library generated by RepeatModeler</p> <p>Wtdbg2_assembly.log: Wtdbg2.5 log file showing exact command used, kmer distribution, memory usage and assembly duration.</p> <p>CAZY_Annotation.tar.gz: dbCAN2 Identification of CAZy in the selected crustacean proteomes as well as list of cellulase-associated GH groups (glycoside hydrolase). </p> <p>Orthofinder.tar.gz: Orthofinder2 output and proteomes of each crustacean used to infer orthologous clustering.</p> <p>GH9_Analysis.tar.gz: Selected GH9-associated protein sequences, amino acid alignment and IQTree output. </p> <p>Cdes_mito.gbf: GenBank file of the annotated complete mitogenome</p> <p>Cdes.filtered.codingseq: Cherax destructor protein-coding genes with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping. </p> <p>Cdes.filtered.faa: Cherax destructor proteins with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping. </p> <p>Cdes.ortholog.list: List of predicted Cherax destructor proteins with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping. </p>
TAGET: A toolkit for analyzing full-length transcripts from long-read sequencing
<p>Polished transcripts of COLO829 from the PacBio platform. The original web link: https://downloads-ap.pacbcloud.com/public/dataset/Melanoma2019_IsoSeq/PolishedMappedTranscripts/before-SQANTI2filter/.</p>
A multiomic characterization of the leukemia cell line REH using short- and long-read sequencing
<p>This is a public repository containing secondary datasets described in the publication <a href="https://doi.org/10.26508/lsa.202302481">"A multiomic characterization of the leukemia cell line REH using short- and long-read sequencing"</a>. Primary data for this project are available at NCBI/SRA under the BioProject accession numbers PRJNA600820 and PRJNA834955, and include the following sequencing datasets:</p> <p>REH cell line:</p> <ul> <li>PacBio WGS</li> <li>ONT Ultralong WGS</li> <li>Illumina short-read PCR-free WGS</li> <li>IsoSeq RNA-seq</li> <li>Illumina short-read RNA-seq</li> </ul> <p>GM12878 cell line:</p> <ul> <li>Illumina short-read RNA-seq</li> </ul> <p>This dataset includes the following files:</p> <p><strong>Depth of Coverage analysis</strong></p> <ul> <li>Output from `samtools coverage`: <em>samtools.coverage.illumina.txt, samtools.coverage.ont.txt, samtools.coverage.pb.txt</em></li> <li>Output from `copycat` (binned coverage): <em>copycat.ont.coverage.10kb.csv, copycat.pb.coverage.10kb.csv, copycat.pcrfree.coverage.10kb.csv</em></li> </ul> <p><strong>Structural Variant (SV) callsets</strong></p> <ul> <li><strong>Raw:</strong> <em>illumina.tiddit.vcf, ont.sniffles.vcf, pb.sniffles.vcf</em></li> <li><strong>Filtered: </strong><em>REH.svs.filtered.csv</em></li> </ul> <p><strong>SNV callsets</strong></p> <ul> <li><strong>Filtered and annotated:</strong> <em>REH.mutect.filtered.ann.vcf.gz</em></li> </ul> <p><strong>Fusion gene callsets</strong></p> <ul> <li><strong>Short-read: </strong><em>GM12878.fusionreport.txt, illumina.all.txt, illumina.filtered.csv, REH.arriba.fusions.tsv, REH.fusioncatcher.fusion-genes.txt, REH.pizzly.txt, REH.squid.fusions.annotated.txt, REH.starfusion.abridged.tsv, REH.pdf</em></li> <li><strong>Long-read: </strong><em>cupcake.long.csv, cupcake.std.csv, jaffa_results.csv</em></li> <li><strong>Filtered: </strong><em>REH.fusions.filtered.csv</em><br> </li> </ul> <p> </p>
A roadmap to durable BCTV resistance using long-read genome assembly of genetic stock KDH13
<p>PacBio Sequence data associated with genetic stock KDH13.Datasets include genome resources and annotated files associated with the manuscript "Long-read genome assembly of Double Haploid Sugar Beet KDH13 provides roadmap for durable genetic resistance to Beet Curly Top Virus". This includes genome assembly, ordered genome assembly, protein predictions, variant call format files for an F1 hybrid (KDH13xKDH19-17).</p>
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
<p><i>Telopea speciosissima, </i>the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for <i>T. speciosissima</i>. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta BUSCOs 'Complete'. We present a new method in Diploidocus (<a href="https://github.com/slimsuite/diploidocus">https://github.com/slimsuite/diploidocus</a>) for classifying, curating and QC-filtering scaffolds, which combines read depths, <i>k</i>-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (<a href="https://github.com/slimsuite/depthsizer">https://github.com/slimsuite/depthsizer</a>), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2<i>n</i> = 22). Genome annotation predicted 40,158<code> </code>protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated <i>CYCLOIDEA </i>(<i>CYC</i>)<i> </i>genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 'Duplicated' BUSCO genes using a new tool, DepthKopy (<a href="https://github.com/slimsuite/depthkopy">https://github.com/slimsuite/depthkopy</a>), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level <i>T. speciosissima</i> reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.</p>
Long-read DNA-Seq of SmAP1 knockout strains
<p>Halobacterium salinarum knockout strains (delta-Ura3: control; delta-Ura3/delta-SmAP1: SmAP1 knockout) were cultured in uracil-supplemented (50 μg/mL) complex medium (CM) until mid-exponential phase (OD600nm = 0.5). The cultures were grown at 37C, under light exposure, and with constant agitation (125 RPM). We collected 2 mL samples and submitted them to DNA extraction using the DNeasy Blood & Tissue kit (QIAGEN), according to the manufacturer's instructions for Gram-negative bacteria. We tested DNA samples for purity and quantity using spectrophotometric and fluorimetric methods, respectively. The samples were prepared for long-read sequencing following the 1D native barcoding genomic DNA protocol using SQK-LSK108 and EXP-NBD103 (Oxford Nanopore Technologies). The equimolar pool of barcoded samples was sequenced using a MinION Mk1B instrument (Oxford Nanopore Technologies) in an FLO-MIN106 flow cell for 24 hours. Three biological replicates were sequenced for each one of the strains (control and SmAP1 knockout). This repository stores the partitioned (eleven parts: aa-ak) compressed directory (tar.gz) containing all the raw reads (fast5 format) output by the MinKNOW software.</p> <p><strong>Experimental design:</strong></p> <table> <tbody> <tr> <td><strong>Barcode</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Barcode 01</td> <td>Control, biological replicate 1</td> </tr> <tr> <td>Barcode 02</td> <td>Control, biological replicate 2</td> </tr> <tr> <td>Barcode 03</td> <td>Control, biological replicate 3</td> </tr> <tr> <td>Barcode 04</td> <td>SmAP1 knockout, biological replicate 1</td> </tr> <tr> <td>Barcode 05</td> <td>SmAP1 knockout, biological replicate 2</td> </tr> <tr> <td>Barcode 06</td> <td>SmAP1 knockout, biological replicate 3</td> </tr> </tbody> </table> <p><strong>Instructions to merge files and extract:</strong></p> <p>1. Download all the files available in this Zenodo entry (smap1_ko_exp_fast5.tar.gz.part_a*; from aa to ak; eleven files) to your preferred directory;</p> <p>2. Execute the following commands using a Linux or OSX terminal:</p> <pre><code class="language-bash"># concatenate all the files into a single one cat smap1_ko_exp_fast5.tar.gz.part_a* > smap1_ko_exp_fast5.tar.gz # extract the merged file tar zxvf smap1_ko_exp_fast5.tar.gz</code></pre> <p> </p>
Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing
<p>Studies in vertebrate genomics require sampling from a broad range of tissue types, taxa, and localities. Recent advancements in long-read and long-range genome sequencing have made it possible to produce high-quality chromosome-level genome assemblies for almost any organism. However, adequate tissue preservation for the requisite ultra-high molecular weight DNA (uHMW DNA) remains a major challenge. Here we present a comparative study of preservation methods for field and laboratory tissue sampling, across vertebrate classes and different tissue types. We find that no single method is best for all cases. Instead, the optimal storage and extraction methods vary by taxa, by tissue, and by down-stream application. Therefore, we provide sample preservation guidelines that ensure sufficient DNA integrity and amount required for use with long-read and long-range sequencing technologies across vertebrates. Our best practices generate the uHMW DNA needed for the high-quality reference genomes for Phase 1 of the Vertebrate Genomes Project (VGP), whose ultimate mission is to generate chromosome-level reference genome assemblies of all ~70,000 extant vertebrate species.</p>
Genome-resolved metagenomics using short-, long-read and metaHiC sequencing
<p>Reference-quality metagenome-assembled genomes (MAGs) are the key to exploring microbial compositions and microbe-phenotype associations. They can be recovered by different sequencing technologies and computational tools, which need an unbiased and comprehensive assessment to identify best practices. This work systematically evaluates 40 distinct strategies to recover high-quality MAGs generated by eight assemblers, eight metagenomics binners, and four sequencing technologies, including short-, long-read and metaHiC sequencing. We notice that the hybrid assemblies of short- and long-reads outperform either short- or long-read assemblies and generate more contigs with high contiguity. When the hybrid assemblies are combined with metaHiC-based binning (Hybrid-HiC), more high-quality MAGs with higher taxonomic diversity are recovered, and more tRNA and rRNA genes, phages, plasmids and antibiotic resistance genes are identified in the mock, simulated and real datasets.</p>
Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2
<p><span>Long-read sequencing technologies have improved significantly since their emergence. Their read lengths, potentially spanning entire transcripts, is advantageous for reconstructing transcriptomes. Existing long-read transcriptome assembly methods are primarily reference-based and to date, there is little focus on reference-free transcriptome assembly. We introduce RNA-Bloom2, a reference-free assembly method for long-read transcriptome sequencing data. </span>RNA-Bloom2 is available on GitHub at: <a href="https://github.com/bcgsc/RNA-Bloom">https://github.com/bcgsc/RNA-Bloom</a>.</p> <p><span>We benchmarked the assembly quality and the computational performance of RNA-Bloom2 on simulated data. We prepared two mouse simulated datasets with Trans-NanoSim</span><span> for the cDNA and dRNA sequencing protocols model on experimental ONT data</span><span>. The datasets were simulated </span><span>based on the mouse ENSEMBL annotation for GRCm39.</span><span> To investigate the effect of sequencing depth, we subsampled each dataset to 2, 10, and 18 million reads, resulting in a total of six sets of reads for our benchmarking experiments. Using the simulated data, w</span><span>e showed that the transcriptome assembly quality of RNA-Bloom2 is competitive to those of reference-based methods.</span></p>
Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Real and mock datasets
<p>This repository contains the reads, assemblies, and references required to replicate the <strong>real and mock</strong> results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>
Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Simulated datasets
<p>This repository contains the reads, assemblies, and references required to replicate the <strong>simulated</strong> results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>
Novel canine high-quality metagenome-assembled genomes by long-read metagenomics together with Hi-C proximity ligation
<p>We characterized a canine fecal sample of a healthy dog by combining a long-read metagenomics assembly (Nanopore sequencing) with Hi-C cross-linking data, and further correction of the frameshift errors. We retrieved and characterized 27 HQ MAGs and seven MQ MAGs considering MIMAG criteria.</p> <p>Find in this repository the final Hi-C genomics bins (CanMAG_XX-HiCbin.fa), including both the genome and the extra-chromosomal elements within the bin. </p> <p> </p>
ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data (repository for simulated ONT RNA-seq data)
<p>Simulated ONT direct RNA and 1D cDNA sequencing data of varying sequencing depths (0.5 million, 1 million, 3 million, and 5 million simulated reads) used for benchmark evaluations of transcript discovery and quantification in our paper "ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data". All details can be found in the <strong>Materials and Methods</strong> section of the paper. </p> <p><em>HEK293T_DirectRNA.transcriptome_quantification.tsv</em> and <em>HEK293T_DirectRNA.transcriptome_quantification.tsv </em>are tab-separated files containing estimated raw read counts and normalized abundance values (in TPM) of transcripts annotated in GENCODE v34lift37. Transcript quantification was done using NanoSim (version 3.1.0). </p> <p><em>HEK293T_DirectRNA.NanoSim_500k.fastq.gz</em>,<em> </em><em>HEK293T_DirectRNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_DirectRNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_DirectRNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT direct RNA sequencing reads respectively. </p> <p><em>HEK293T_1DcDNA.NanoSim_500k.fastq.gz</em>,<em> HEK293T_1DcDNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_1DcDNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_1DcDNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT 1D cDNA sequencing reads respectively. </p>
Multiplexed long-read plasmid validation and analysis using OnRamp
<p>Plasmid read data and references from "Multiplexed long-read plasmid validation and analysis using OnRamp "</p> <p>Experiments:</p> <ol> <li>AAZ605 - 7 plasmids</li> <li>AFQ178 - 9 plasmids</li> <li>ACK577 - 30 plasmids</li> <li>AEZ576 - 15 plasmids</li> <li>plasmids_ref_7.fasta</li> <li>plasmids_ref_9.fasta</li> <li>plasmids_ref_15.fasta</li> <li>plasmids_ref_30.fasta</li> </ol> <p> </p>
Long-read sequencing of diagnosis and post-therapy medulloblastoma reveals complex rearrangement patterns and epigenetic signatures
<p>Imaging data related to the "Long-read sequencing of diagnosis and post-therapy medulloblastoma reveals complex rearrangement patterns and epigenetic signatures" manuscript</p>
Fibertools: fast and accurate m6A calling using single-molecule long-read sequencing (ML data)
<p>Fibertools is a convolutional neural network that permits the fast and accurate identification of endogenous and exogenous N6-methyladenine (m6A)-marked bases using single-molecule long-read sequencing.<strong> </strong>This dataset (ML data) provides training and validation data for training fibertools supervised and semi-supervised CNN models for three long-read chemistries.</p>
A high-quality, long-read genome assembly of the whitelined sphinx moth (Lepidoptera: Sphingidae: Hyles lineata)
<p><span>The sphinx moth genus <em>Hyles</em> comprises 29 described species inhabiting all continents except Antarctica. The genus diverged relatively recently (40 – 25 mya), arising in the Americas and rapidly establishing a cosmopolitan distribution. The whitelined sphinx moth, <em>Hyles lineata</em>, represents the oldest extant lineage of this group and is one of the most widespread and abundant sphinx moths in North America. <em>Hyles lineata </em>exhibits the large body size and adept flight control characteristic of the sphinx moth family (Sphingidae), but is unique in displaying extreme larval color variation and broad host plant use. These traits, in combination with its broad distribution and high relative abundance within its range, have made <em>H. lineata</em> a strong model organism for studying phenotypic plasticity, plant-herbivore interactions, physiological ecology, and flight control. Despite being one of the most well-studied sphinx moths, little data exists on genetic variation or regulation of gene expression. Here we report a high-quality draft genome showing high contiguity (N50 of 14.2 Mb) and completeness (98.2% of Lepidoptera BUSCO genes), an important first characterization to facilitate such studies. We also annotate the core melanin synthesis pathway genes and confirm that they have high sequence conservation with other moths and are most similar to those of another, well-characterized sphinx moth, the tobacco hornworm (<em>Manduca sexta</em>).</span></p>
Metadata for: Environmental adaptations by the intertidal Antarctic cyanobacterium Halotia branconii CENA392 as revealed using long-read genome sequencing
<p>Antarctica poses numerous challenges to life such as cold shock, low nutrient concentrations and periodic desiccation over a wide range of extreme temperatures. Cyanobacteria survive this harsh environment having evolved adaptive metabolic plasticity to become the dominant primary producers. The type strain cyanobacterium <em>Halotia branconii</em> CENA392 was isolated from an Antarctic intertidal seashore. The complete circular genome of this strain is presented herein, which was assembled using long sequence reads. The genome encoded some stress-related genes associated with low-temperature adaptation and biosynthesis of mycosporine-like amino acid (MAA) photoprotective compounds. Empirical experimentation demonstrated constitutive production of the MAA porphyra-334 and total carotenoids without exposure to low temperatures or ultraviolet radiation stress. Phylogenetic analysis provided insights on the taxonomic placement and the evolutionary history of some annotated genes. These data exemplify the importance of generating complete quality genome sequences of microorganisms isolated from extreme intertidal environments, facilitating in-depth evaluation of ecological and taxonomic inferences.</p>
HG002 data for Profiling Chromatin Accessibility in Humans Using Adenine Methylation and Long-Read Sequencing
<p>This dataset includes 6mA frequency data for the HG002 native DNA (untreated) sample sequenced on nanopore r9.4.1.</p>
NA12878 and MCF7 data for Profiling Chromatin Accessibility in Humans Using Adenine Methylation and Long-Read Sequencing
<p>This dataset includes 5mC and 6mA frequency data for NA12878 and MCF7 EcoGII-treated chromatin samples sequenced on nanopore r9.4.1.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.