Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

525

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

525 results for “Long read”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset for "Progessive improvement of the Australian blacklip abalone (Haliotis rubra) genome assembly with Nanopore long reads, hybrid meta assembly and haplotig purging

<p>This Zenodo archive contains the blacklip abalone genome assemblies and and their BUSCO completeness calculations. Genome annotation (gff3 format), CDS, protein sequences and Orthofinder2 output were also included.</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Polished Assemblies for "GoldPolish-Target: Targeted long-read genome assembly polishing"

<p>GoldPolish-Target is a targeted genome assembly polishing tool that uses long reads. We tested GoldPolish-Target on Oxford Nanopore Technologies datasets with a human cell line (NA24385) and Drosophila melanogaster (fruit fly). Here, we provide the data for the GoldRush baseline (unpolished) assembly and the GoldPolish-Target and medaka polished assemblies of these long-read datasets.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

TEST DATA for Enhanced protein isoform characterization through long-read proteogenomics

<p>Test data for&nbsp;The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a></li> <li><a href="http://10.5281/zenodo.5920920">Long-Read-Proteogenomics Workflow Results using Jurkat Sample data</a></li> </ol> <p>This Repository contains the test data, specifically:</p> <p><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Baseline assemblies for "ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads" protocol

<p>ntLink is a flexible&nbsp;<em>de novo</em> genome scaffolding toolkit which can be run in various modes depending on the desired user output, with multiple new functionalities recently introduced. Here, we provide the baseline assembly datasets used in the ntLink protocol paper &quot;ntLink: a toolkit for <em>de novo </em>genome assembly scaffolding and mapping using long reads&quot;. The provided assemblies are ABySS (short-read) and Flye (long-read) assemblies of&nbsp;<em>Caenorhabditis elegans&nbsp;</em>genome sequencing data. The ABySS (v2.1.4) assembly utilized paired-end short reads (accession&nbsp;DRR008444), and was run with the following parameters:&nbsp;k=64&nbsp;l=40 s=1000&nbsp;q=15 B=10G j=8&nbsp;kc=3&nbsp;H=4 S=1000-10000 N=9.The&nbsp;<em>C. elegans</em>&nbsp;Flye (v2.5) assembly was run using Oxford Nanopore long reads (accession SRR10028109) and the following parameters:&nbsp;--nano-raw SRR10028109.fastq&nbsp;-g100m -t48.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

De novo assembly of a long-read Amblyomma americanum genome

<p>Genome assemblies of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the unphased diploid assembly generated by the Flye assembler (Arcadia_Amblyomma_americanum_asm001.fasta). In addition, there are two associated fasta files containing sequences generated by submitting the unphased diploid assembly to separation by the Purge_Dups pipeline&nbsp;(purged pseudo-haploid assembly and haplotig assembly).</p> <p>Flye assembler:&nbsp;https://github.com/fenderglass/Flye</p> <p>Purge_Dups pipeline:&nbsp;https://github.com/dfguan/purge_dups</p> <p>NCBI Bioproject:&nbsp;PRJNA932813</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

De novo assembly of a long-read Amblyomma americanum genome (NCBI/Genbank deposited genome)

<p>Genome assembly&nbsp;of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the phased pseudo-haploid tick genome&nbsp;generated after assembly using&nbsp;Flye, phasing using&nbsp;Purge_Dups, and clean-up using custom python scripts generated in-house.&nbsp;</p> <p>NCBI Bioproject:&nbsp;PRJNA932813</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Assemblies for "Linear time complexity de novo long read genome assembly with GoldRush"

<p>GoldRush is&nbsp;a&nbsp;<em>de novo</em>&nbsp;genome assembly algorithm with linear time complexity in the number of input long sequencing reads. We tested GoldRush on Oxford Nanopore Technologies datasets with different base error profiles describing the genomes of three human cell lines (NA24385, HG01243 and HG02055),&nbsp;Oryza sativa&nbsp;(rice), and&nbsp;Solanum lycopersicum&nbsp;(tomato). Here, we provide the assemblies for the GoldRush, Flye, Redbean and Shasta assemblies of these long read datasets.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Long-read, chromosome-scale assembly of Vitis rotundifolia cv. Carlos and its unique resistance to Xylella fastidiosa subsp. fastidiosa.

<p>We assembled and annotated a new, long-read genome assembly for &lsquo;Carlos&rsquo;, a cultivar of muscadine that exhibits tolerance, to build upon the existing genetic resources available for muscadine. We are awaiting release of the genome through NCBI, so we have made the assembly and annotations available here.</p>

opencc-by-4.0May 2023View details →
dryad40/100

Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads

<p>In the paper associated with this dataset, a workflow is presented which enables the identification of single-/low-copy nuclear molecular markers for a plant group of interest, by mining data from a small representative target capture experiment done using a commercial probe kit and Oxford Nanopore long-read sequencing. The proposed pipeline first assesses sequence variability contained in the data from targeted loci and assigns reads to their respective genes, via a combined BLAST/clustering procedure. Cluster consensus sequences are then examined based on four pre-defined criteria presumably indicative for absence of paralogy. This is done by calculating four specialized indices; loci are ranked according to their performance in these indices, and top-scoring loci are considered putatively single- or low-copy. The approach can be applied to any probe set. As it relies on long reads, the contribution also provides template workflows for processing Nanopore-based target capture data. Identified loci can be used for NGS amplicon sequencing. For detection of possibly remaining paralogy in these data, which might occur in groups with rampant paralogy, the long-read assembly tool CANU is employed. The presented workflow can be useful for researchers dealing with reticulate or polyploidization phylogenetic histories in plants.</p> <p>The present dataset contains several documents supplementing the original paper. Its most important elements are a detailed description (alongside two graphical workflow figures) of all methods employed in the study, suitable for reproducing the steps of the workflow and also the wet-lab work. The workflow employs a collection of BASH, Python and R scripts which is available here, together with a detailed account on command line use in Linux. Also, reference sequences for the identified markers can be found as well as sequence alignments derived from the amplicon sequencing.</p>

opencc-zeroJun 2023View details →
dryad40/100

Data from: Improved genome assembly of the whiteleg shrimp Penaeus (Litopenaeus) vannamei using long- and short-read sequences from public databases

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad40/100

Long- and short-read metabarcoding technologies reveal similar spatio-temporal structures in fungal communities

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad40/100

Reconstructing NOD-like receptor alleles with high internal conservation in Podospora anserina using long-read sequencing

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad40/100

A novel method to assess the integrity of frozen archival DNA samples: Alpha-diversity ratios of short and long-read 16S rRNA gene sequences

Open the record for dataset details and reuse information.

publicAug 2024View details →
zenodo36/100

Raw Nanopore data for "Nanopore Long-Read Guided Complete Genome Assembly of Hydrogenophaga intermedia, and Genomic Insights into 4-Aminobenzenesulfonate, p-Aminobenzoic Acid and Hydrogen Metabolism in the Genus Hydrogenophaga"

<p>This is the raw Nanopore dataset (fast5) for Hydrogenophaga intermedia PBC. The gDNA was prepared using the now obsolete SQK-NSK007 kit and sequenced on a MINION R9 Flowcell.&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Dataset for "Nanopore-led long-read genome assembly of the Australian yabby, Cherax destructor"

<p>Intermediate_Assemblies.tar.gz: Intermediate genome assemblies e.g. raw wtdbg assembly&nbsp;(CD.raw.fa), polished wtdbg assembly&nbsp;(CD.cns.fa), 1st pilon polished assembly (CDF2_pilon1.fasta), 2nd pilon polished assembly (CDF2_pilon2.fasta) and RNA-scaffolded assembly (CDF2_pilon2_prna.fasta). Folders with run_&quot;assembly name&quot; are BUSCO output for each of the assembly.</p> <p>BRAKER2.tar.gz:&nbsp;BRAKER2 genome annotation output containing the initial set of predicted protein-coding genes as well as&nbsp;training intermediate files.</p> <p>BUSCOv3.tar.gz: BUSCO assessment of publicly available&nbsp;Decapod crustacean genome assemblies</p> <p>Cdes.filtered.codingseq: Filtered&nbsp;set of protein-coding sequences</p> <p>Cdes.filtered.faa: Translation of the filtered&nbsp;protein-coding sequences</p> <p>CDF2.NCBI.fasta.masked.gz: Repeat-masked (softmasked)&nbsp;Cherax destructor genome</p> <p>Cqua_transcriptome.tar.gz: rnaSPAdes output (combined fasta) of all Cherax quadricarinatus transcriptomes and its reduced dataset generated by EvidentialGene.&nbsp;</p> <p>Quast.tar.gz: Quast output of all Decapod crustacean genome assemblies assessed in this study</p> <p>Repeat_Annotation.tar.gz: Repeat annotation (.gff3) based on Cherax destructor-specific de novo repeat library and its summary (.tbl)</p> <p>RepeatLibrary.tar.gz: Cherax destructor-specific de novo repeat library generated by RepeatModeler</p> <p>Wtdbg2_assembly.log: Wtdbg2.5 log file showing exact command used, kmer distribution, memory usage and assembly duration.</p> <p>CAZY_Annotation.tar.gz: dbCAN2 Identification of CAZy in the selected crustacean proteomes as well as list of cellulase-associated GH groups (glycoside hydrolase).&nbsp;</p> <p>Orthofinder.tar.gz: Orthofinder2 output and proteomes of each crustacean used to infer orthologous clustering.</p> <p>GH9_Analysis.tar.gz: Selected GH9-associated protein sequences, amino acid alignment and IQTree output.&nbsp;</p> <p>Cdes_mito.gbf: GenBank file of the annotated complete mitogenome</p> <p>Cdes.filtered.codingseq: Cherax destructor protein-coding genes&nbsp;with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping.&nbsp;</p> <p>Cdes.filtered.faa:&nbsp;Cherax destructor proteins with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping.&nbsp;</p> <p>Cdes.ortholog.list: List of predicted Cherax destructor proteins with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping.&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Short and Long Read Sequences from VIM-1 producing (Klebsiella pneumoniae test-dataset)

<p>Creation of a downsized version of the long and short reads data of VIM-1 producing <i>Klebsiella pneumoniae</i>. Samples were obtained from the Hospital Infanta Cristina (Badajoz, Spain) on October 11th, 2019.&nbsp;</p><p>- Long read data: Oxford Nanopore Technology (A1403KPN.fastq.gz).</p><p>- Short read data : Illumina (A1403KPN_R2_filtered.fastq.gz &amp;&amp; A1403KPN_R2_filtered.fastq.gz)</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

TAGET: A toolkit for analyzing full-length transcripts from long-read sequencing

<p>Polished transcripts&nbsp;of COLO829 from the PacBio platform. The original web link:&nbsp;https://downloads-ap.pacbcloud.com/public/dataset/Melanoma2019_IsoSeq/PolishedMappedTranscripts/before-SQANTI2filter/.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

De novo genome assembly of rice varieties using Nanopore long reads

<p>Genome sequences for Sugimura et al. (2024) of the rice (O. sativa) varieties 'Hitomebore' and 'Arroz da Terra.'</p> <p>Yusaku Sugimura, Kaori Oikawa, Yu Sugihara, Hiroe Utsushi, Eiko Kanzaki, Kazue Ito, Yumiko Ogasawara, Tomoaki Fujioka, Hiroki Takagi, Motoki Shimizu, Hiroyuki Shimono, Ryohei Terauchi, Akira Abe. Impact of rice GENERAL REGULATORY FACTOR14h (GF14h) on low-temperature seed germination and its application to breeding. PLoS Genet 20(8): e1011369. https://doi.org/10.1371/journal.pgen.1011369</p> <p>bioRxiv doi: https://doi.org/10.1101/2024.02.16.580620</p>

opencc-by-4.0Jan 2024View details →
dryad36/100

Long read genome assembly of Automeris io (Lepidoptera: Saturniidae) an emerging model for the evolution of deimatic displays

<p>Automeris moths are a morphologically diverse group with 145 described species that have a geographic range that spans from the New World temperate zone to the Neotropics. Many Automeris have hindwing eyespots that are thought to deter or disrupt the attack of potential predators, allowing the moth time to escape. Some species in the genus have vestigial eyespots or lack them completely, suggesting that this trait may provide a selective benefit. The Io moth (Automeris io), known for its striking eyespots, is the most widely studied species within the genus and is an emerging model system to study the evolution of deimatism, a predatory defense that combines visual stimuli and movement. Here we present a high-quality, PacBio HiFi genome assembly for Io moth to aid existing research on the molecular development of eyespots. Genomic research is needed to address questions involving antipredatory defenses and eyespot pattern development. BUSCO analysis for this genome shows a completeness of 98.4%, and N50 of 15.</p>

opencc-zeroFeb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record