Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

233

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

233 results for “long-read”

Learn how ShareScore rates datasets ↗
zenodo36/100

metaFlye: scalable long-read metagenome assembly using repeat graphs

<p>Supplementary files for the manuscript titled &quot;metaFlye: scalable long-read metagenome assembly using repeat graphs&quot;.</p> <p>The archive includes generated assemblies and the corresponding metaQUAST evaluations.</p>

opencc-by-4.0May 2019View details →
dryad36/100

Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad36/100

A high-quality, long-read genome assembly of the whitelined sphinx moth (Lepidoptera: Sphingidae: Hyles lineata)

Open the record for dataset details and reuse information.

publicApr 2023View details →
dryad36/100

Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad36/100

A roadmap to durable BCTV resistance using long-read genome assembly of genetic stock KDH13

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad36/100

Metadata for: Environmental adaptations by the intertidal Antarctic cyanobacterium Halotia branconii CENA392 as revealed using long-read genome sequencing

Open the record for dataset details and reuse information.

publicMay 2023View details →
zenodo32/100

NanoGalaxy: Nanopore long-read sequencing data analysis in Galaxy

<p>The data presented in &quot;NanoGalaxy: A Galaxy tool kit with workflows for third-generation sequence analysis&quot; to illustrate the functionality of the tools was obtained from: Wick, Ryan R., et al. &quot;Completing bacterial genome assemblies with multiplex MinION sequencing.&quot;&nbsp;<em>Microbial genomics</em>&nbsp;3.10 (2017).</p> <p>+</p> <p>Li, Ruichao, et al. &quot;Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data.&quot;&nbsp;<em>Gigascience</em>&nbsp;7.3 (2018): gix132.</p>

opencc-by-4.0Apr 2020View details →
dryad32/100

Construction of a chromosome-scale long-read reference genome assembly for potato

<p><span><strong>Background:</strong> Worldwide, the cultivated potato, <i>Solanum tuberosum </i>L<i>.</i>, is the number one vegetable crop and a critical food security crop. The genome sequence of DM1-3 516 R44, a doubled monoploid clone of S. <i>tuberosum </i>Group Phureja, was published in 2011 using a whole-genome shotgun sequencing approach with short read sequence data. Current advanced sequencing technologies now permit generation of near-complete, high-quality chromosome-scale genome assemblies at a minimal cost. </span></p> <p><span><strong>Findings: </strong>Here, we present an updated version of the DM1-3 516 R44 genome sequence (v6.1) using Oxford Nanopore Technologies long reads coupled with proximity-by-ligation scaffolding (Hi-C) yielding a chromosome-scale assembly. The new (v6.1) assembly represents 741.6 Mb of sequence (87.8 %) of the estimated 844 Mb genome, of which, 741.5 Mb is non-gapped with 731.2 Mb anchored to the 12 chromosomes. Use of Oxford Nanopore Technologies full-length cDNA sequencing enabled annotation of 32,917 high-confidence protein-coding genes encoding 44,851 gene models that had a significantly improved representation of conserved orthologs compared to the previous annotation. The new assembly has improved  contiguity with a 595-fold increase in N50 contig size, 99% reduction in the numbersof contigs, a 44-fold increase in N50 scaffold size, and an LTR Assembly Index score of 13.56, placing it in the category of reference genome quality. The improved assembly also permitted annotation of the centromeres via alignment to sequencing reads derived from CENH3 nucleosomes. </span></p> <p><span><strong>Conclusions: </strong>Access to advanced sequencing technologies and improved software permitted generation of a high-quality, long-read, chromosome-scale assembly and improved annotation dataset for the reference genotype of potato that will facilitate research aimed at improving agronomic traits and understanding genome evolution.</span></p>

opencc-zeroAug 2020View details →
dryad32/100

Improved contiguity of the threespine stickleback genome using long-read sequencing

<p>While the cost and time for assembling a genome has drastically decreased, it still remains a challenge to assemble a highly contiguous genome. These challenges are rapidly being overcome by the integration of long-read sequencing technologies. Here, we use long-read sequencing to improve the contiguity of the threespine stickleback fish (Gasterosteus aculeatus) genome, a prominent genetic model species. Using Pacific Biosciences sequencing, we assembled a highly contiguous genome of a freshwater fish from Paxton Lake. Using contigs from this genome, we were able to fill over 76% of the gaps in the existing reference genome assembly, improving contiguity over five-fold. Our gap filling approach was highly accurate, validated by 10X Genomics long-distance linked-reads. In addition to closing a majority of gaps, we were able to assemble segments of telomeres and centromeres throughout the genome. This highlights the power of using long sequencing reads to assemble highly repetitive and difficult to assemble regions of genomes. This latest genome build has been released through a newly designed community genome browser that aims to consolidate the growing number of genomics datasets available for the threespine stickleback fish.</p>

opencc-zeroJan 2021View details →
zenodo32/100

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing

<p>These are the VCF files of structural variant (SV) calls for two sibling patients (II:2 [GMPB009_1] and II:3 [GMPB009_4]) generated by PacBio HiFi long-read genome sequencing.</p> <p>Sequence reads were processed using the <a href="https://github.com/PacificBiosciences/pb-human-wgs-workflow-snakemake">PacBio Human WGS workflow</a> with the human reference genome (hg38), and SVs were identified using '<a href="https://github.com/PacificBiosciences/svpack">svpack</a>'.</p>

opencc-by-4.0Apr 2024View details →
dryad32/100

Long-read genome sequencing of bread wheat facilitates disease resistance gene cloning

<p>Cloning agronomically important genes from large, complex crop genomes remains challenging. Here, we generate a 14.7-gigabase chromosome-scale<i> </i>assembly of the South African bread wheat (<i>Triticum aestivum</i>) cultivar Kariega by combining high-fidelity long reads, optical mapping, and chromosome conformation capture. The resulting assembly is an order of magnitude more contiguous than previous wheat assemblies. Kariega shows durable resistance against the devastating fungal stripe rust disease. We identified the race-specific disease resistance gene <i>Yr27</i>, encoding an intracellular immune receptor, as a major contributor to this resistance. <i>Yr27</i> is allelic to the leaf rust resistance gene <i>Lr13,</i> with the Yr27 and Lr13 proteins sharing 97% sequence identity. Our results thus demonstrate the feasibility of generating chromosome-scale wheat assemblies to clone genes and also exemplify that highly similar alleles of a single-copy gene can confer resistance to different pathogens, which might provide a basis for engineering <i>Yr27</i> alleles with multiple recognition specificities in future.</p>

opencc-zeroDec 2021View details →
zenodo32/100

Supplementary material 1 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431

Table S1. List of specimens used for this study, including their voucher information, plate location, indexing, amplicon concentration and sequencing results, both as an output from SMRT tools (CCSs) and as an output from DADA2 (sequence variants). Table S2. List of the 64 barcode sequences used to index the samples. Used barcode pairs are listed in Table S1

opencc-zeroFeb 2022View details →
zenodo32/100

TPM quantification (StringTie) of long-read isoforms in TCGA breast tumors and GTEx

<p>StringTie was applied to quantify the expression of long-read isoforms sequenced in TCGA and GTEx samples.</p> <p>Inputs:</p> <ul> <li>RNA-seq samples (bam files) from breast cancer samples from TCGA (tumors) and GTEx tissues (healthy controls).</li> <li>GTF: full-length isoforms sequenced using PacBio long-read sequencing in breast tumors. See&nbsp;<a href="https://www.science.org/doi/10.1126/sciadv.abg6711">A comprehensive long-read isoform analysis platform and sequencing resource for breast cancer</a>.</li> </ul> <p>Output: tab-delimited files with TPM values for isoforms parsed from StringTie results.</p> <p>More details:&nbsp;<a href="https://www.science.org/doi/10.1126/sciadv.abg6711">A comprehensive long-read isoform analysis platform and sequencing resource for breast cancer</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Supplement table of mesozooplankton taxa obtained using long-read and short-read metabarcoding

<p>Supplement tables containing information about publications on mesozooplankton taxa in the Ross Sea using two metabarcoding analyses</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

Methods for structural variant detection with long-read sequencing data

<p>SV calls from different long-read based SV callers on sequencing data. SV results evaluated in&nbsp;Methods for structural variant detection with long-read sequencing data.</p> <p>NA24385_Pacbio_HiFi -&gt; HiFi_L1 in paper</p> <p>NA24385_Pacbio_MtSinai -&gt; CLR_L1 in paper</p> <p>NA24385_Pacbio_CLR_SRX7668835 -&gt; CLR_L2 in paper</p> <p>NA24385_Pacbio_CLR_SRX6719924 -&gt; CLR_L3 in paper</p> <p>NA24385_ONT_Promethion -&gt; Nano_L1 in paper</p>

opencc-by-4.0Oct 2022View details →
dryad32/100

Data from: Haplotype sequence collection of ABO blood group alleles by long-read sequencing reveals putative A1-diagnostic variants

<p>In the era of blood group genomics, reference collections of complete and fully-resolved blood group gene alleles have gained high importance. For most blood groups, however, such collections are currently lacking, as resolving full-length gene sequences as haplotypes (i.e. separated maternal/paternal origin) remains exceedingly difficult with both Sanger and short-read next-generation sequencing. Using the latest third-generation long-read sequencing, we generated a collection of fully-resolved sequences for all six main <em>ABO</em> allele groups: <em>ABO</em>*<em>A1</em>/<em>A2</em>/<em>B</em>/<em>O.01.01</em>/<em>O.01.02</em>/<em>O.02</em>. We selected 77 samples from an <em>ABO</em> genotype dataset (n=25,200) of serologically-typed Swiss blood donors. The entire <em>ABO</em> gene was amplified in two overlapping long-range PCRs (covering ~23.6 kb) and sequenced by long-read Oxford Nanopore sequencing. For quality validation, two samples per <em>ABO</em> group were re-sequenced using Illumina and PacBio technology. All 154 full-length <em>ABO</em> sequences were resolved as haplotypes. We observed novel, distinct sequence patterns for each <em>ABO</em> group. Most genetic diversity was found between, not within, <em>ABO</em> groups. Phylogenetic tree and haplotype network analyses highlighted distinct clades of each <em>ABO</em> group. Strikingly, our data uncovered four genetic variants putatively specific for <em>ABO</em>*<em>A1</em>, for which direct diagnostic targets are currently lacking. We validated <em>A1</em>-diagnostic potential using whole-genome data (n=4,872) of a multi-ethnic cohort. Overall, our sequencing strategy proved powerful for producing high-quality <em>ABO</em> haplotypes and holds promise for generating similar collections for other blood groups. The publicly available collection of 154 haplotypes will serve as a valuable resource for molecular analyses of <em>ABO</em>, as well as studies about the function and evolutionary history of <em>ABO</em>.</p>

opencc-zeroOct 2022View details →
zenodo32/100

the transcriptome GTFs, FASTA and SQANTI reports for short-read assembled isoforms, long-read assembled isoforms and our assembled isoforms

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

AsaruSim: a single-cell and spatial RNA-Seq Nanopore long-reads simulation workflow

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo32/100

Comprehensive Structural Variant Benchmark Dataset: 1100 VCF files from long-read sequencing of 10 NCBI individuals

<p>We initially collected 10 NCBI individuals: HG002 family pedigree data (HG002 [son], HG003 [father], HG004 [mother]), the HG005 family pedigree data (HG005 [son], HG006 [father], HG007 [mother]), the NA12878 subject, the HG00096 subject, the HG00512 subject and the CHM13 subject. Then we used PacBio (CLR: Continuous Long Read, CCS: Circular Consensus Sequencing) and Nanopore (ONT) platforms, 5 aligners and 10 callers to construct the pipelines, with most parameters set to default values. After that, except for 6 invalid pipelines(pbmm2-Nanovar, lra-Picky, lra-delly, lra-NanoVar, lra-NanoSV, lra-pbsv), we obtain 1100 VCF files.</p>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record