Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
233
datasets available to search
ShareScore release 0.9.0
Dataset results
233 results for “long-read”
metaFlye: scalable long-read metagenome assembly using repeat graphs
<p>Supplementary files for the manuscript titled "metaFlye: scalable long-read metagenome assembly using repeat graphs".</p> <p>The archive includes generated assemblies and the corresponding metaQUAST evaluations.</p>
Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing
Open the record for dataset details and reuse information.
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
Open the record for dataset details and reuse information.
A high-quality, long-read genome assembly of the whitelined sphinx moth (Lepidoptera: Sphingidae: Hyles lineata)
Open the record for dataset details and reuse information.
Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2
Open the record for dataset details and reuse information.
A roadmap to durable BCTV resistance using long-read genome assembly of genetic stock KDH13
Open the record for dataset details and reuse information.
Metadata for: Environmental adaptations by the intertidal Antarctic cyanobacterium Halotia branconii CENA392 as revealed using long-read genome sequencing
Open the record for dataset details and reuse information.
NanoGalaxy: Nanopore long-read sequencing data analysis in Galaxy
<p>The data presented in "NanoGalaxy: A Galaxy tool kit with workflows for third-generation sequence analysis" to illustrate the functionality of the tools was obtained from: Wick, Ryan R., et al. "Completing bacterial genome assemblies with multiplex MinION sequencing." <em>Microbial genomics</em> 3.10 (2017).</p> <p>+</p> <p>Li, Ruichao, et al. "Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data." <em>Gigascience</em> 7.3 (2018): gix132.</p>
Construction of a chromosome-scale long-read reference genome assembly for potato
<p><span><strong>Background:</strong> Worldwide, the cultivated potato, <i>Solanum tuberosum </i>L<i>.</i>, is the number one vegetable crop and a critical food security crop. The genome sequence of DM1-3 516 R44, a doubled monoploid clone of S. <i>tuberosum </i>Group Phureja, was published in 2011 using a whole-genome shotgun sequencing approach with short read sequence data. Current advanced sequencing technologies now permit generation of near-complete, high-quality chromosome-scale genome assemblies at a minimal cost. </span></p> <p><span><strong>Findings: </strong>Here, we present an updated version of the DM1-3 516 R44 genome sequence (v6.1) using Oxford Nanopore Technologies long reads coupled with proximity-by-ligation scaffolding (Hi-C) yielding a chromosome-scale assembly. The new (v6.1) assembly represents 741.6 Mb of sequence (87.8 %) of the estimated 844 Mb genome, of which, 741.5 Mb is non-gapped with 731.2 Mb anchored to the 12 chromosomes. Use of Oxford Nanopore Technologies full-length cDNA sequencing enabled annotation of 32,917 high-confidence protein-coding genes encoding 44,851 gene models that had a significantly improved representation of conserved orthologs compared to the previous annotation. The new assembly has improved contiguity with a 595-fold increase in N50 contig size, 99% reduction in the numbersof contigs, a 44-fold increase in N50 scaffold size, and an LTR Assembly Index score of 13.56, placing it in the category of reference genome quality. The improved assembly also permitted annotation of the centromeres via alignment to sequencing reads derived from CENH3 nucleosomes. </span></p> <p><span><strong>Conclusions: </strong>Access to advanced sequencing technologies and improved software permitted generation of a high-quality, long-read, chromosome-scale assembly and improved annotation dataset for the reference genotype of potato that will facilitate research aimed at improving agronomic traits and understanding genome evolution.</span></p>
Improved contiguity of the threespine stickleback genome using long-read sequencing
<p>While the cost and time for assembling a genome has drastically decreased, it still remains a challenge to assemble a highly contiguous genome. These challenges are rapidly being overcome by the integration of long-read sequencing technologies. Here, we use long-read sequencing to improve the contiguity of the threespine stickleback fish (Gasterosteus aculeatus) genome, a prominent genetic model species. Using Pacific Biosciences sequencing, we assembled a highly contiguous genome of a freshwater fish from Paxton Lake. Using contigs from this genome, we were able to fill over 76% of the gaps in the existing reference genome assembly, improving contiguity over five-fold. Our gap filling approach was highly accurate, validated by 10X Genomics long-distance linked-reads. In addition to closing a majority of gaps, we were able to assemble segments of telomeres and centromeres throughout the genome. This highlights the power of using long sequencing reads to assemble highly repetitive and difficult to assemble regions of genomes. This latest genome build has been released through a newly designed community genome browser that aims to consolidate the growing number of genomics datasets available for the threespine stickleback fish.</p>
Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing
<p>These are the VCF files of structural variant (SV) calls for two sibling patients (II:2 [GMPB009_1] and II:3 [GMPB009_4]) generated by PacBio HiFi long-read genome sequencing.</p> <p>Sequence reads were processed using the <a href="https://github.com/PacificBiosciences/pb-human-wgs-workflow-snakemake">PacBio Human WGS workflow</a> with the human reference genome (hg38), and SVs were identified using '<a href="https://github.com/PacificBiosciences/svpack">svpack</a>'.</p>
Long-read genome sequencing of bread wheat facilitates disease resistance gene cloning
<p>Cloning agronomically important genes from large, complex crop genomes remains challenging. Here, we generate a 14.7-gigabase chromosome-scale<i> </i>assembly of the South African bread wheat (<i>Triticum aestivum</i>) cultivar Kariega by combining high-fidelity long reads, optical mapping, and chromosome conformation capture. The resulting assembly is an order of magnitude more contiguous than previous wheat assemblies. Kariega shows durable resistance against the devastating fungal stripe rust disease. We identified the race-specific disease resistance gene <i>Yr27</i>, encoding an intracellular immune receptor, as a major contributor to this resistance. <i>Yr27</i> is allelic to the leaf rust resistance gene <i>Lr13,</i> with the Yr27 and Lr13 proteins sharing 97% sequence identity. Our results thus demonstrate the feasibility of generating chromosome-scale wheat assemblies to clone genes and also exemplify that highly similar alleles of a single-copy gene can confer resistance to different pathogens, which might provide a basis for engineering <i>Yr27</i> alleles with multiple recognition specificities in future.</p>
Supplementary material 1 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431
Table S1. List of specimens used for this study, including their voucher information, plate location, indexing, amplicon concentration and sequencing results, both as an output from SMRT tools (CCSs) and as an output from DADA2 (sequence variants). Table S2. List of the 64 barcode sequences used to index the samples. Used barcode pairs are listed in Table S1
TPM quantification (StringTie) of long-read isoforms in TCGA breast tumors and GTEx
<p>StringTie was applied to quantify the expression of long-read isoforms sequenced in TCGA and GTEx samples.</p> <p>Inputs:</p> <ul> <li>RNA-seq samples (bam files) from breast cancer samples from TCGA (tumors) and GTEx tissues (healthy controls).</li> <li>GTF: full-length isoforms sequenced using PacBio long-read sequencing in breast tumors. See <a href="https://www.science.org/doi/10.1126/sciadv.abg6711">A comprehensive long-read isoform analysis platform and sequencing resource for breast cancer</a>.</li> </ul> <p>Output: tab-delimited files with TPM values for isoforms parsed from StringTie results.</p> <p>More details: <a href="https://www.science.org/doi/10.1126/sciadv.abg6711">A comprehensive long-read isoform analysis platform and sequencing resource for breast cancer</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
Supplement table of mesozooplankton taxa obtained using long-read and short-read metabarcoding
<p>Supplement tables containing information about publications on mesozooplankton taxa in the Ross Sea using two metabarcoding analyses</p>
Methods for structural variant detection with long-read sequencing data
<p>SV calls from different long-read based SV callers on sequencing data. SV results evaluated in Methods for structural variant detection with long-read sequencing data.</p> <p>NA24385_Pacbio_HiFi -> HiFi_L1 in paper</p> <p>NA24385_Pacbio_MtSinai -> CLR_L1 in paper</p> <p>NA24385_Pacbio_CLR_SRX7668835 -> CLR_L2 in paper</p> <p>NA24385_Pacbio_CLR_SRX6719924 -> CLR_L3 in paper</p> <p>NA24385_ONT_Promethion -> Nano_L1 in paper</p>
Data from: Haplotype sequence collection of ABO blood group alleles by long-read sequencing reveals putative A1-diagnostic variants
<p>In the era of blood group genomics, reference collections of complete and fully-resolved blood group gene alleles have gained high importance. For most blood groups, however, such collections are currently lacking, as resolving full-length gene sequences as haplotypes (i.e. separated maternal/paternal origin) remains exceedingly difficult with both Sanger and short-read next-generation sequencing. Using the latest third-generation long-read sequencing, we generated a collection of fully-resolved sequences for all six main <em>ABO</em> allele groups: <em>ABO</em>*<em>A1</em>/<em>A2</em>/<em>B</em>/<em>O.01.01</em>/<em>O.01.02</em>/<em>O.02</em>. We selected 77 samples from an <em>ABO</em> genotype dataset (n=25,200) of serologically-typed Swiss blood donors. The entire <em>ABO</em> gene was amplified in two overlapping long-range PCRs (covering ~23.6 kb) and sequenced by long-read Oxford Nanopore sequencing. For quality validation, two samples per <em>ABO</em> group were re-sequenced using Illumina and PacBio technology. All 154 full-length <em>ABO</em> sequences were resolved as haplotypes. We observed novel, distinct sequence patterns for each <em>ABO</em> group. Most genetic diversity was found between, not within, <em>ABO</em> groups. Phylogenetic tree and haplotype network analyses highlighted distinct clades of each <em>ABO</em> group. Strikingly, our data uncovered four genetic variants putatively specific for <em>ABO</em>*<em>A1</em>, for which direct diagnostic targets are currently lacking. We validated <em>A1</em>-diagnostic potential using whole-genome data (n=4,872) of a multi-ethnic cohort. Overall, our sequencing strategy proved powerful for producing high-quality <em>ABO</em> haplotypes and holds promise for generating similar collections for other blood groups. The publicly available collection of 154 haplotypes will serve as a valuable resource for molecular analyses of <em>ABO</em>, as well as studies about the function and evolutionary history of <em>ABO</em>.</p>
the transcriptome GTFs, FASTA and SQANTI reports for short-read assembled isoforms, long-read assembled isoforms and our assembled isoforms
Open the record for dataset details and reuse information.
AsaruSim: a single-cell and spatial RNA-Seq Nanopore long-reads simulation workflow
Open the record for dataset details and reuse information.
Comprehensive Structural Variant Benchmark Dataset: 1100 VCF files from long-read sequencing of 10 NCBI individuals
<p>We initially collected 10 NCBI individuals: HG002 family pedigree data (HG002 [son], HG003 [father], HG004 [mother]), the HG005 family pedigree data (HG005 [son], HG006 [father], HG007 [mother]), the NA12878 subject, the HG00096 subject, the HG00512 subject and the CHM13 subject. Then we used PacBio (CLR: Continuous Long Read, CCS: Circular Consensus Sequencing) and Nanopore (ONT) platforms, 5 aligners and 10 callers to construct the pipelines, with most parameters set to default values. After that, except for 6 invalid pipelines(pbmm2-Nanovar, lra-Picky, lra-delly, lra-NanoVar, lra-NanoSV, lra-pbsv), we obtain 1100 VCF files.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.